Infrastructure and MLOps Engineer

  • Full Time
  • Bristol

Graphcore

Infrastructure and MLOps Engineer Overview

Company Name Graphcore
Job Role Infrastructure and MLOps Engineer
Qualifications Not Specified
Category IT Jobs
Job Type Full Time
Location Bristol

At Graphcore, we are at the forefront of AI compute innovation, combining expertise in semiconductors, software, and AI to create a comprehensive AI compute stack. As part of the SoftBank Group, we are well-funded and focused on delivering cutting-edge technology within the expanding SoftBank AI ecosystem. We are currently looking to expand our teams globally to tackle the exciting challenges in AI.

We invite you to join our dynamic Software Infrastructure team, where you will play a crucial role in scaling and managing our infrastructure. Your work will involve developing essential tools and services that empower our software teams, enhancing the processes of building, testing, deploying, and productizing our Machine Learning Software components. You will gain valuable experience working with our High-Performance Computing (HPC) AI platforms and distributed systems.

The Team

The Software Infrastructure team is responsible for providing critical platforms and services that support software development across the organization. Our duties include managing the Continuous Integration (CI) platform, build engineering, component integration, and packaging and release systems. We operate in squads, promoting a culture of service ownership and empowerment among our engineers, focusing on long-term engineering solutions and minimizing repetitive tasks.

Responsibilities and Duties

  • Develop, own, and maintain tools and services to support AI research and engineering teams.
  • Deploy and maintain services using Kubernetes and Docker.
  • Manage our cloud infrastructure utilizing tools such as Terraform.
  • Enhance the build, test, deployment, and productization processes for Machine Learning Software components.

Candidate Profile

Essential Qualifications:

  • Knowledge of Python programming.
  • Familiarity with cloud services, especially AWS.
  • Experience in managing or developing in Linux environments.
  • Understanding of CI/CD principles.
  • Experience using Kubernetes (k8s).
  • Experience in one of the following areas: maintaining machine learning applications, deploying ML orchestration tools (e.g., NV Ray, KFP, SkyPilot), or managing ML accelerator hardware (e.g., DCGM).

Desirable Qualifications:

  • Experience with Infrastructure as Code (IaC) tools such as Terraform/OpenTofu.
  • Familiarity with GitHub Actions.
  • Experience with modern observability tools like Prometheus and Grafana.
  • Knowledge of programming languages such as Go, Java, or C++.

Benefits

  • Competitive salary.
  • Flexible working arrangements.
  • Generous annual leave policy.
  • Private medical insurance and health cash plan.
  • Dental plan.
  • Pension scheme with matched contributions up to 5%.
  • Life assurance and income protection.
  • Generous parental leave policy.
  • Employee assistance program for health, mental wellbeing, and bereavement support.
  • Access to healthy food and snacks at our central Bristol office, along with a barista bar.

We are committed to fostering an inclusive work environment that welcomes individuals from diverse backgrounds and experiences. We offer an equal opportunity recruitment process and are open to discussing reasonable adjustments during the interview process.


Degree Requirement: Not Specified

Visa Sponsorship May be

To apply for this job please visit job-boards.greenhouse.io.