Performance & Observability Engineer

Performance & Observability Engineer

Full-Time 63000 - 77000 £ / year (est.) No working from home possible
E

At a Glance

  • Tasks: Ensure system reliability and visibility through performance engineering and observability.
  • Company: Join a leading global law firm with a commitment to innovation and excellence.
  • Benefits: Enjoy a competitive salary, professional development, and a diverse, inclusive culture.
  • Other info: Be part of a dynamic team focused on growth and collaboration.
  • Why this job: Make a real impact by optimising technology and enhancing client service.
  • Qualifications: Experience in observability tools and strong problem-solving skills required.

The predicted salary is between 63000 - 77000 £ per year.

Herbert Smith Freehills Kramer is a world-leading global law firm, where our ambition is to help you achieve your goals.

Exceptional client service and the pursuit of excellence are at our core.

We invest in and care about our client relationships, which is why so many are longstanding.

We We enjoy breaking new ground, as we have for over 170 years.

As a fully integrated transatlantic and transpacific firm, we are where you need us to be.

Our footprint is extensive and committed across the world’s largest markets, key financial centres and major growth hubs.

At our best tackling complexity and navigating change, we work alongside you on demanding litigation, exacting regulatory work and complex public and private market transactions.

We are recognised as leading in these areas.

We are immersed in the sectors and challenges that impact you.

We are recognised as standing apart in energy, infrastructure and resources.

And we’re focused on areas of growth that affect every business across the world.

All of this is achieved by supporting the growth of our people, who help us deliver on our ambition – which is to help you achieve yours.

  • Herbert
  • Smith
  • Freehills

Kramer: Your goals.

Our ambition

The Opportunity

The Performance & Observability Engineer plays a critical role in ensuring system reliability, scalability, and visibility across the entire technology stack.

This role focuses on transitioning from traditional monitoring to full observability, enabling deep performance insights, real-time issue detection, and proactive optimisation.

The ideal candidate will have expertise in performance engineering, distributed tracing, logging, metrics, and automation, driving improvements across cloud, infrastructure, and application layers.

  • Primary Responsibilities
  • Performance Engineering & Optimisation Analyse application, database, and infrastructure performance to identify bottlenecks and inefficiencies.
  • Develop performance benchmarks and SLIs to measure service responsiveness and stability.
  • Collaborate with SRE and Dev Ops teams to optimise CI/CD pipelines for performance improvements.
  • Collaborate with wider IT teams to Implement caching strategies, query optimisation, and autoscaling to enhance system efficiency.
  • Observability Platform Development & Implementation
  • Design and implement end-to-end observability frameworks covering metrics, logs, traces, and events.
  • Instrument services using existing tools (eg: Nexthink) to improve visibility.
  • Enable distributed tracing across microservices to enhance root cause analysis and performance debugging.
  • Standardise logging and telemetry collection across infrastructure, applications, and cloud services.
  • Define best practices and consistent approach across development teams to improve monitoring consistency.
  • Maturing from Monitoring to Full Observability Transition from basic alerting to proactive insights, leveraging AI-driven anomaly detection.
  • Ensure comprehensive observability across frontend, backend, databases, cloud infrastructure, and networking.
  • Implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to track system health.
  • Automate root cause analysis and incident detection through advanced monitoring techniques.
  • Incident Response & Reliability Engineering
  • Reduce Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) through improved observability.
  • Integrate monitoring and alerting tools with incident response platforms (ie: Service Now).
  • Develop self-healing and auto-remediation mechanisms to reduce operational toil.
  • Improve alerting strategies by reducing false positives and improving signal-to-noise ratio.
  • Governance, Compliance & Reporting
  • Define observability best practices and governance models to ensure adoption across teams.
  • Ensure log retention, security, and compliance with standards (e. g., GDPR, SOC 2, PCI DSS).
  • Develop executive dashboards and reporting frameworks to showcase reliability and performance trends.
  • Key Performance Indicators
  • Maturity of Observability Capabilities % of Services with Full Observability Coverage – Ensure visibility across the entire stack.
  • Instrumentation Completeness (%) – Track the number of services fully instrumented with logs, metrics, and traces.
  • Service-Level Indicator (SLI) Coverage – Ensure key performance indicators are defined and tracked.
  • Reduction in Blind Spots (%) – Improve monitoring coverage across all components.
  • Performance & Reliability Metrics Application Response Time (P99, P95, P50 Latency) – Improve service performance.
  • System Throughput & Load Handling (%) – Increase service efficiency and scalability.
  • Reduction in Performance Bottlenecks (%) – Optimise infrastructure and application layers.
  • Successful Load Test Pass Rate (%) – Ensure applications meet expected performance benchmarks.
  • Incident Management & Operational Efficiency Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) – Reduce system downtime and recovery time.
  • Reduction in Noisy Alerts (%) – Improve alert relevance and reduce false positives.
  • Proactive Issue Detection (%) – Increase the percentage of incidents identified before user impact.
  • Error Budget Utilization (%) – Ensure system reliability is balanced with innovation velocity.
  • Automation & AI-Driven Observability (AIOps) % of Issues Resolved via Automated Remediation – Reduce manual intervention in incident response.
  • Reduction in On-Call Burden (%) – Minimize alerts requiring human intervention.
  • Anomaly Detection Accuracy (%) – Improve proactive detection of performance issues.
  • Business & Customer Impact
  • Customer Experience Metrics (Latency, Errors, Uptime) – Ensure observability drives tangible business improvements.
  • Downtime Reduction (%) – Improve service availability and reliability.
  • Cost Optimisation from Performance & Observability (%) – Reduce operational inefficiencies and cloud expenses.

Qualifications, Skills and Experience

Technical Skills

Experience in using and maintaining Observability & APM Tools – Grafana experience is essential Experience of using KQL Performance Testing & Load Testing Cloud & Infrastructure Monitoring – AWS Cloud Watch, Azure Monitor, GCP Operations Suite, Kubernetes Observability.

Knowledge of Log Aggregation & Analysis – eg: Grafana Loki, Splunk etc Automation & Scripting – Power Automate, Terraform, Power Shell Incident Response & ITSM – Service Now.

Sound understanding of firm’s applications, systems and tools across technology stack Dev Ops experience is desirable Nexthink experience is desirable

Soft Skills & Collaboration

Strong problem-solving and root cause analysis skills.

Ability to translate observability insights into business impact for stakeholders.

A continuous improvement mindset, focused on reducing toil and improving efficiency.

Experience working in a Dev Ops, SRE, or Platform Engineering environment.

  • Team
  • Information Technology
  • Working Pattern
  • Full time
  • Location
  • London
  • Contract
  • Permanent
  • Diversity & Inclusion

We are committed to attracting people from all backgrounds and creating a respectful and inclusive culture where everyone thrives.

We see this as essential to our success, including our ability to innovate and achieve sustained high performance.

This is a key part of our Values—Human, Bold, and Outstanding.

At Herbert Smith Freehills Kramer, we align your growth and ambition with ours.

We invest in your personal and professional growth and support you to achieve your ambitions.

And you share responsibility for playing a part in delivering the firm’s growth and ambition too.

A leading global law firm, with over 6,000 people, we are in the world\'s largest markets, key financial centres and major growth hubs.

We’re recognised leaders in demanding contentious matters, exacting regulatory work and complex public and private transactions.

We’re immersed in the many challenges facing our clients.

We’re invested and resourceful.

We understand the part technology and digitalisation play in the delivery of legal services.

We want to make a positive impact wherever in the world we operate.

You’ll have the opportunity to engage with this with an open mind and curiosity.

We are known for our diverse perspectives and renowned for our culture.

Being human, bold and outstanding are more than our values: you’ll discover they are our lived experience.

And by being ambitious for your growth and ours, we’ll achieve our goals together.

  • Herbert
  • Smith
  • Freehills

Kramer: Your growth.

Our ambition.

#J-18808-Ljbffr

Performance & Observability Engineer employer: Exchange House Services Ltd - London

Herbert Smith Freehills Kramer is an exceptional employer, offering a dynamic work environment in the heart of London where you can thrive as a Senior Associate in Corporate Private Equity. With a strong commitment to diversity and inclusion, the firm fosters a culture that values innovation and collaboration, providing ample opportunities for professional development and career progression. Employees benefit from engaging in high-profile transactions while being supported in their personal growth, making it a rewarding place to build a meaningful career.

E

Contact Details:

Exchange House Services Ltd - London Recruitment Team

We think you need these skills to ace Performance & Observability Engineer

Performance Engineering
Observability Framework Development
Distributed Tracing
Logging and Metrics Analysis
Automation and Scripting
Cloud Infrastructure Monitoring
Incident Response