AI Harness & Evaluations Engineers
Capgemini
AI Harness & Evaluations Engineers Overview
| Company Name | Capgemini |
| Job Role | AI Harness & Evaluations Engineers |
| Qualifications | Bachelor’s |
| Category | IT Jobs |
| Job Type | Full Time |
| Location | London |
Capgemini is seeking experienced AI Harness & Evaluations Engineers to join their innovative team. This role involves managing the evidence that demonstrates the trustworthiness of AI products within highly regulated sectors such as finance and insurance. You will be responsible for developing and maintaining high-quality datasets in collaboration with domain experts, especially in claims and payments, where correctness is based on professional judgment rather than simple labels.
Your work will include overseeing the operation of evaluation harnesses, designing evaluation gates like thresholds and calibration procedures, and integrating these processes into continuous integration pipelines as merge gates and into production environments as drift monitors that automatically score live traffic samples. You will also develop systems to analyze production failures, transforming them into regression cases to prevent recurrence, and prepare evaluation evidence packages that meet external model-risk validation standards, ensuring they are understandable to non-engineering stakeholders.
This role requires hands-on experience with machine learning or data engineering, with a focus on rigorous measurement practices and working with production systems that have been validated through measurement. Strong statistical knowledge, including cohort design, sample sizing, and confidence intervals, is essential. You should have practical experience in evaluation engineering for language models, including creating golden datasets, calibrating models as judges, and implementing regression gates within CI systems. Proficiency in Python and familiarity with CI/CD systems are also necessary.
Candidates should be willing to develop domain expertise by working closely with claims or payments professionals daily, and have experience constructing adversarial datasets, conducting human annotation programs, and understanding safety testing boundaries. Knowledge of reinforcement learning pipelines and evaluation methodologies is advantageous.
The ideal applicant will have experience with relevant technology stacks such as self-hosted LangSmith and LangGraph platforms, PostgreSQL with pgvector, ClickHouse, S3-compatible storage, Neo4j Enterprise, MCP-native connectors, OpenTelemetry, Grafana, and Kubernetes with Helm and Argo CD, or similar systems. The role is based in London with a hybrid working pattern, combining office, client site, and home working, tailored to business needs.
Capgemini emphasizes a collaborative, quality-focused environment where engineers write specifications, harnesses, evaluations, and guardrails, with AI agents executing implementation loops under human review at key stages. The organization values diversity and inclusion, being a Level 2 Disability Confident Employer, and offers a comprehensive benefits package aimed at employee wellbeing and professional growth. Successful candidates will contribute to delivering trustworthy AI solutions that meet regulatory standards and client expectations, with a focus on continuous improvement and innovation.
Degree Requirement: Bachelor’s
Visa Sponsorship May be
To apply for this job please visit careers.capgemini.com.