Confirmed on the employer's own hiring board on Oct 9, 2026. First seen by Alion on Sep 6, 2026.
Join our team as a Senior Machine Learning Engineer, where you'll be responsible for building the judgement layer of our agent evaluation platform. This role involves developing rubrics, judges, and calibration methodologies, as well as fine-tuning small models and evaluating their performance. You'll work on applied machine learning in a dynamic and challenging environment, with opportunities for growth and impact.
Missions
- Construire la couche de jugement de notre plateforme d'évaluation des agents, y compris les rubriques, les juges, la calibration par rapport aux étiquettes humaines.
- Développer un juge calibré qui peut évaluer une trajectoire et servir de signal de récompense pour l'optimisation de l'agent.
- Collaborer avec l'équipe d'annotation pour établir une boucle de calibration continue contre les trajectoires étiquetées par des humains.
Profil recherché
- Fine-tuning and evaluating small models: SFT, preference tuning, distillation- 5+ years in applied ML, data science, or ML-adjacent engineering, with a track record of work that shipped and got used
- Strong Python, and the discipline to ship production-grade code rather than notebooks
- Experience in at least 3 of these:
- Search ranking, recsys, or online experimentation evaluation - golden-set staleness, offline/online divergence, side-by-side rater agreement. This is the closest existing analog to agentic eval, and it transfers directly
- Ability to think and communicate clearly about complex problems - a large part of this job is convincing engineers that a number means what you say it means, and being right
- Reward modeling, RLHF/RLAIF, or process reward models
- A high degree of ownership and a bias toward shipping at startup pace
- Prompt engineering as an engineering discipline - versioned, tested, and measured, not tuned by vibes
- Strong applied ML fundamentals, and comfort treating LLMs as a component you evaluate, prompt, and fine-tune rather than one you pretrain
- Experience turning subjective human judgement into a measurement that holds up - one that other people, and ideally other models, can act on. This is the core of the job
- Human annotation programs: rubric authoring, label quality, and annotator throughput as a real constraint
- Agent trajectory analysis and step-level fault attribution
- LLM-as-judge or automated evaluation design, and calibrating it against human judgement
- Comfort with ambiguity, and the judgement to know when a measurement is good enough to act on

