ML Ops Engineer

Anaplan

ML Ops Engineer Overview

Company Name Anaplan
Job Role ML Ops Engineer
Qualifications Not Specified
Category IT Jobs
Job Type Full Time
Location London

At Anaplan, we are seeking a talented Machine Learning Operations (ML Ops) Engineer to join our Platform Engineering team in London. Our mission is to develop and sustain scalable, high-performance infrastructure that supports our innovative AI-driven scenario planning platform. This role involves working closely with Data Scientists, ML Engineers, and Cloud Infrastructure specialists to streamline the processes of training, deploying, and running inference on machine learning models, with a focus on maximizing GPU utilization, system reliability, and cost efficiency.

Key Responsibilities

  • Design, develop, and maintain cloud-native infrastructure for MLOps and LLMOps, ensuring it supports the needs of our AI-infused platform.
  • Collaborate with cross-functional teams to optimize model training, deployment, and inference workflows, focusing on performance and cost-effectiveness.
  • Manage containerized environments using Kubernetes and Docker, along with GPU orchestration frameworks such as NVIDIA GPU Operator, Slurm, or Ray.
  • Automate infrastructure provisioning and management through Infrastructure as Code tools like Terraform, Helm, and Ansible.
  • Enhance GPU workload performance, networking, and storage solutions to facilitate efficient training and low-latency inference.
  • Build and sustain CI/CD pipelines tailored for machine learning models, enabling continuous training, evaluation, and deployment.
  • Deploy Large Language Models and generative AI workloads using inference engines like Triton Inference Server, vLLM, or TensorRT-LLM.
  • Implement automated validation, monitoring for model and data drift, and latency issues to ensure ongoing system performance.
  • Monitor and optimize cloud spending across AWS, GCP, and Azure, implementing auto-scaling, spot instances, and dynamic resource management to reduce costs.
  • Establish benchmarking and telemetry systems to measure throughput and economics of training and inference operations.
  • Set up comprehensive observability using tools such as Prometheus, Grafana, OpenTelemetry, and MLflow or Weights & Biases.

Required Skills and Experience

  • Practical experience in DevOps, SRE, or Platform Engineering, especially related to AI/ML infrastructure in production settings.
  • Proven track record of deploying, scaling, and operationalizing machine learning models and LLMs in cloud environments.
  • Experience managing GPU-intensive compute environments and high-performance computing (HPC) systems.
  • Expertise in Kubernetes, Docker, Helm, Kubeflow, and service meshes like Istio.
  • Hands-on experience with Infrastructure as Code tools such as Terraform and Ansible, and CI/CD pipelines using GitHub Actions, ArgoCD, or Jenkins.
  • Familiarity with inference frameworks including vLLM, Ray, MLflow, DeepSpeed, or Hugging Face TGI.
  • Strong knowledge of cloud platforms such as AWS, GCP, and Azure, with skills in GPU cost optimization techniques and tools like Kubecost.
  • Proficiency in programming languages including Python, Bash, or Go, and a deep understanding of Linux system tuning and performance monitoring.

Our Commitment

We are dedicated to fostering a diverse, equitable, inclusive, and welcoming environment. We believe that embracing different perspectives and backgrounds enhances our innovation and success. We are committed to providing reasonable accommodations for individuals with disabilities during the application and employment process. Our hiring practices are designed to respect and value every individual, regardless of gender, ethnicity, age, neurodiversity, or other protected characteristics.

Beware of fraudulent job offers. Anaplan does not send unsolicited job offers via email or phone without an extensive interview process. All official communications will come from an @anaplan.com email address. If you have doubts about any communication, contact us at [email protected].

If you are interested in building your career with us, you can create a job alert to receive future opportunities directly to your email. To apply, please fill out the application form with your details, including your legal and preferred names, contact information, and work authorization status.


Degree Requirement: Not Specified

Visa Sponsorship May be

To apply for this job please visit job-boards.greenhouse.io.