Infrastructure and MLOps Engineer
Graphcore
Infrastructure and MLOps Engineer Overview
| Company Name | Graphcore |
| Job Role | Infrastructure and MLOps Engineer |
| Qualifications | Not Specified |
| Category | IT Jobs |
| Job Type | Full Time |
| Location | Cambridge |
At Graphcore, we are at the forefront of AI compute, combining expertise in semiconductors, software, and artificial intelligence to create a comprehensive AI compute stack. As part of the SoftBank Group, we are well-funded and positioned to deliver crucial technology within the expanding SoftBank AI ecosystem. We are looking to grow our teams globally to tackle significant challenges and make a meaningful impact in the field of artificial intelligence.
We are currently seeking an Infrastructure and MLOps Engineer to join our Software Infrastructure team. This role is essential for scaling and managing our infrastructure, where you will develop vital tools and services that enhance the productivity of our software teams. Your work will directly contribute to improving the build, test, deployment, and productization processes of our Machine Learning Software components, while also providing you with valuable experience in High-Performance Computing (HPC) AI platforms and distributed systems.
Team Overview
The Software Infrastructure team plays a crucial role in delivering platforms and services that support software development across the organization. Our responsibilities include managing the Continuous Integration (CI) platform, build engineering, component integration, and packaging and release systems. We operate in squads, promoting a culture of service ownership and empowerment among our engineers, with a focus on long-term engineering solutions and minimizing repetitive tasks.
Key Responsibilities
- Develop, own, and maintain tools and services to support AI research and engineering teams.
- Deploy and maintain services using Kubernetes and Docker.
- Manage our Cloud Infrastructure with tools such as Terraform.
Candidate Profile
We are looking for candidates with the following qualifications:
Essential Skills
- Proficiency in Python.
- Familiarity with cloud services, especially AWS.
- Experience in managing or developing in Linux environments.
- Understanding of CI/CD principles.
- Experience with Kubernetes (k8s).
- Experience in one of the following: maintaining machine learning applications, deploying ML orchestration tools (e.g., NV Ray, KFP, SkyPilot), or managing ML accelerator hardware (e.g., DCGM).
Desirable Skills
- Experience with Infrastructure as Code (IaC) tools such as Terraform/OpenTofu.
- Familiarity with GitHub Actions.
- Experience with modern observability tools like Prometheus and Grafana.
- Knowledge of programming languages such as Go, Java, or C++.
Benefits
- Competitive salary.
- Flexible working arrangements.
- Generous annual leave policy.
- Private medical insurance and health cash plan.
- Dental plan.
- Pension scheme with a 5% match.
- Life assurance and income protection.
- Generous parental leave policy.
- Employee assistance program for health, mental wellbeing, and bereavement support.
- Access to healthy food and snacks at our office, including a barista bar.
We are committed to creating an inclusive work environment and welcome individuals from diverse backgrounds. We encourage applicants to discuss any reasonable adjustments needed during the interview process.
Degree Requirement: Not Specified
Visa Sponsorship May be
To apply for this job please visit job-boards.greenhouse.io.