Lead Site Reliability Engineer

  • Full Time
  • UK

Nice

Lead Site Reliability Engineer Overview

Company Name Nice
Job Role Lead Site Reliability Engineer
Qualifications Not Specified
Category IT Jobs
Job Type Full Time
Location Rest of UK

Join our team at NICE Public Safety, where we provide cutting-edge solutions for the Public Safety & Justice market. As a Lead Site Reliability Engineer, you will play a crucial role in ensuring our cloud platforms are observable, measurable, reliable, scalable, and maintainable. This hands-on position is essential as we expand our Cloud Platform Engineering team to continue delivering exceptional service to our global customer base.

Key Responsibilities

  • Collaborate with a team of Site Reliability Engineers to manage production systems and improve reliability.
  • Lead investigations into outages, performance issues, and cost-related challenges.
  • Drive automation initiatives for repetitive tasks while balancing project delivery requirements.
  • Provide technical guidance to Cloud Operations and Support teams, overseeing the products and services they manage.
  • Work with DevOps and engineering teams to set and enforce Service Level Objectives (SLOs), Service Level Agreements (SLAs), and error budgets.
  • Develop and configure monitoring dashboards and alerts using tools such as Grafana and Azure Monitor.
  • Install and configure observability platforms, including Grafana, Prometheus, Azure Monitor, and OpenTelemetry.
  • Create Bicep modules for monitoring infrastructure and deploy them effectively.
  • Regularly review and optimize system performance, cost, and security.

Qualifications

  • A minimum of 6 years of experience in Site Reliability Engineering or related fields.
  • Strong technical, analytical, and troubleshooting abilities.
  • In-depth knowledge of databases and data handling formats such as MS-SQL, Elasticsearch, YML, JSON, and XML.
  • Experience with Azure cloud services.
  • Proficiency in programming or advanced scripting languages, including Python, PowerShell, and C#.
  • Significant experience with infrastructure as code and version control systems, particularly ARM, BICEP, and Git.
  • Expertise in managing monitoring, alerting, and dashboarding platforms like Azure Monitor, Prometheus, Grafana, and Elasticsearch.
  • Demonstrated experience in supporting live cloud services and platforms.
  • Familiarity with Kubernetes and containerization technologies, especially Azure Kubernetes Service (AKS).
  • Exposure to Azure DevOps pipelines for continuous integration and delivery (CI/CD) is preferred.
  • Strong understanding of cyber security principles and compliance frameworks.
  • Excellent communication skills, with the ability to engage effectively with both customers and internal teams.

Benefits

  • Opportunity to work in a dynamic and innovative environment.
  • Access to career growth and development opportunities.
  • Participation in employee events and a supportive workplace culture.
  • Hybrid work model allowing flexibility in work arrangements.

Degree Requirement: Not Specified

Visa Sponsorship May be

To apply for this job please visit boards.eu.greenhouse.io.