Lead Site Reliability Engineer
Nice
Lead Site Reliability Engineer Overview
| Company Name | Nice |
| Job Role | Lead Site Reliability Engineer |
| Qualifications | Not Specified |
| Category | IT Jobs |
| Job Type | Full Time |
| Location | Rest of UK |
Join our team at NICE Public Safety, where we provide cutting-edge solutions for the Public Safety & Justice market. As a Lead Site Reliability Engineer, you will play a crucial role in ensuring our cloud platforms are observable, measurable, reliable, scalable, and maintainable. This hands-on position is essential as we expand our Cloud Platform Engineering team to continue delivering exceptional service to our global customer base.
Key Responsibilities
- Collaborate with a team of Site Reliability Engineers to manage production systems and improve reliability.
- Lead investigations into outages, performance issues, and cost-related challenges.
- Drive automation initiatives for repetitive tasks while balancing project delivery requirements.
- Provide technical guidance to Cloud Operations and Support teams, overseeing the products and services they manage.
- Work with DevOps and engineering teams to set and enforce Service Level Objectives (SLOs), Service Level Agreements (SLAs), and error budgets.
- Develop and configure monitoring dashboards and alerts using tools such as Grafana and Azure Monitor.
- Install and configure observability platforms, including Grafana, Prometheus, Azure Monitor, and OpenTelemetry.
- Create Bicep modules for monitoring infrastructure and deploy them effectively.
- Regularly review and optimize system performance, cost, and security.
Qualifications
- A minimum of 6 years of experience in Site Reliability Engineering or related fields.
- Strong technical, analytical, and troubleshooting abilities.
- In-depth knowledge of databases and data handling formats such as MS-SQL, Elasticsearch, YML, JSON, and XML.
- Experience with Azure cloud services.
- Proficiency in programming or advanced scripting languages, including Python, PowerShell, and C#.
- Significant experience with infrastructure as code and version control systems, particularly ARM, BICEP, and Git.
- Expertise in managing monitoring, alerting, and dashboarding platforms like Azure Monitor, Prometheus, Grafana, and Elasticsearch.
- Demonstrated experience in supporting live cloud services and platforms.
- Familiarity with Kubernetes and containerization technologies, especially Azure Kubernetes Service (AKS).
- Exposure to Azure DevOps pipelines for continuous integration and delivery (CI/CD) is preferred.
- Strong understanding of cyber security principles and compliance frameworks.
- Excellent communication skills, with the ability to engage effectively with both customers and internal teams.
Benefits
- Opportunity to work in a dynamic and innovative environment.
- Access to career growth and development opportunities.
- Participation in employee events and a supportive workplace culture.
- Hybrid work model allowing flexibility in work arrangements.
Degree Requirement: Not Specified
Visa Sponsorship May be
To apply for this job please visit boards.eu.greenhouse.io.