Confirmed on the employer's own hiring board on Oct 8, 2026. First seen by Alion on Sep 4, 2026. Π scores B on the Alion truth index.
Join our ML Infrastructure team as a Machine Learning Infrastructure Engineer. You will be responsible for designing, implementing, and maintaining systems for large-scale model training, optimizing performance, and enabling rapid iteration. You will work closely with researchers to scale JAX-based training across TPU and GPU clusters and contribute to core training code. Strong software engineering fundamentals and experience in ML training infrastructure are required.
Missions
- Conception, mise en œuvre et maintenance des systèmes pour l'entraînement de modèles à grande échelle, y compris la planification, la gestion des tâches, le point de contrôle et la journalisation.
- Collaboration avec les chercheurs pour étendre l'entraînement basé sur JAX sur des clusters TPU et GPU, en minimisant les frictions.
- Profilage et amélioration de l'utilisation de la mémoire, de l'utilisation des appareils, du débit et de la synchronisation distribuée.
Profil recherché
- Strong software engineering fundamentals and experience building ML training infrastructure or internal platforms- Strong cross-functional communication and ownership mindset
- Experience managing training workloads on cloud platforms (e.g., SLURM, Kubernetes, GCP TPU/GKE, AWS)
- Familiarity with distributed training, multi-host setups, data loaders, and evaluation pipelines
- Ability to debug and optimize performance bottlenecks across the training stack
- Hands-on large-scale training experience in JAX (preferred), PyTorch
- Experience designing abstractions that balance researcher flexibility with system reliability
- Background in robotics, multimodal models, or large-scale foundation models
- Experience operating close to hardware (GPU/TPU performance tuning)
- Deep ML systems background (e.g., training compilers, runtime optimization, custom kernels)

