NLP Performance Engineer
G-Research
NLP Performance Engineer Overview
| Company Name | G-Research |
| Job Role | NLP Performance Engineer |
| Qualifications | Bachelor’s |
| Category | IT Jobs |
| Job Type | Full Time |
| Location | London |
G-Research is seeking an exceptional NLP Performance Engineer to join our NLP Engineering team, focusing on the performance of large-scale LLM inference. This role is integral to our engineering team, where you will collaborate with researchers to maximize the performance of LLMs and shape the tools and infrastructure that support our NLP capabilities.
Key Responsibilities
- Profile, benchmark, and optimize large-scale LLM inference workloads across our compute infrastructure.
- Ensure efficient deployment of the latest models across various GPU architectures, adapting the inference stack as hardware evolves.
- Design and implement inference optimizations while maintaining output quality.
- Develop reference implementations, libraries, and tooling to enhance the efficiency and reliability of NLP workloads.
- Collaborate with researchers, senior stakeholders, and engineers to design optimized solutions.
- Work with systems, architecture, and platform teams to evolve the compute stack and influence long-term platform decisions.
Who Are We Looking For?
The ideal candidate will possess a blend of deep knowledge in LLM inference and strong software engineering skills, with a scientific approach to performance. Candidates should have:
- A BachelorĂ¢??s, MasterĂ¢??s, or PhD in computer science, or equivalent experience.
- Proven experience in profiling, benchmarking, and optimizing large-scale LLM inference workloads.
- A scientific, evidence-led approach to performance optimization, utilizing rigorous benchmarking and reproducible measurement.
- Deep understanding of transformer inference, including prefill versus decode, KV-cache behavior, attention variants, and performance bottlenecks.
- Hands-on experience with LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM, or TGI, as well as the PyTorch ecosystem.
- Experience with inference optimization techniques, including quantization, speculative decoding, and model parallelism across modern GPU architectures.
- Strong software engineering skills, particularly in Python and CUDA, and experience in building reliable systems for machine learning workloads.
- Excellent communication skills, with the ability to collaborate effectively across research, infrastructure, and engineering teams.
Why Join Us?
We offer a highly competitive compensation package, including an annual discretionary bonus, and a range of benefits designed to support our employees. These include:
- Lunch provided through Just Eat for Business and access to a dedicated barista bar.
- 35 days of annual leave.
- 9% company pension contributions.
- Comprehensive healthcare and life assurance.
- Cycle-to-work scheme.
- Monthly company events.
- Relocation support covering costs associated with moving and immigration, including a concierge service and expert immigration guidance.
If you are ready to take the next step in your career and make a significant impact in the field of NLP, we encourage you to apply.
Degree Requirement: Bachelor’s
Visa Sponsorship Promising
To apply for this job please visit www.gresearch.com.