The role
We are looking for an early-career engineer who enjoys understanding how modern AI systems actually work.
You will work across model experimentation, inference, evaluation, deployment, and ML systems performance. This is a hands-on engineering role: you will build prototypes, run experiments, investigate failures, profile systems, and turn promising ideas into working implementations.
What you will do
- Experiment with LLMs and multimodal models.
- Build prototypes and production-quality components in Python and PyTorch.
- Run and optimise model inference using tools such as vLLM.
- Measure and improve latency, throughput, batching, caching, GPU utilisation, and memory usage.
- Build evaluation frameworks to understand model quality and behaviour.
- Work with embeddings, retrieval, reranking, structured generation, and tool-using/agentic systems.
- Deploy model-backed services and troubleshoot them in realistic workloads.
- Read relevant research papers and reproduce or test promising ideas.
- Design controlled experiments and analyse failures across data, model, software, and infrastructure.
- Document findings clearly, including what worked, what failed, and why.
Core requirements
- Strong Python skills.
- Working knowledge of PyTorch and modern neural-network architectures.
- Understanding of transformers, tokenisation, embeddings, attention, sampling, and decoding.
- Ability to write maintainable software beyond notebooks.
- Comfortable working with Linux, Git, Docker, APIs, and basic cloud infrastructure.
Strong analytical and debugging skills.
Useful ML systems knowledge
You should understand, or be motivated to learn:
- model serving and inference;
- vLLM or similar runtimes;
- continuous batching and KV caching;
- quantization;
- throughput vs latency trade-offs;
- GPU memory constraints;
- mixed precision and device placement;
- profiling and out-of-memory debugging;
- structured/constrained generation.
Experience with CUDA, Triton, distributed systems, Kubernetes, NCCL, or low-level optimisation is useful but not required.
What we look for
We care more about demonstrated technical depth than years of experience.
Good evidence includes:
- a substantial ML or systems project;
- research or thesis work;
- reproducing or implementing a research paper;
- open-source contributions;
- building or profiling an inference/training system;
- technically serious side projects;
- benchmarks or experiments where you measured and improved performance.
You should be able to explain what you built, why you built it that way, what you measured, what failed, and what you learned.
Academic background
A strong foundation in a quantitative discipline such as Computer Science, Mathematics, Statistics, Engineering, Physics, Operations Research, or a related field is preferred.
Research experience is useful but not mandatory.

