You will...
- Build standardized distributed training frameworks for research and production, drive our training towards new levels of stability and efficiency.
- Comprehensively profile model runtime and memory to pinpoint performance bottlenecks.
- Identify and evaluate emerging technologies that can be adopted into Waabi’s training and inference frameworks. Examples include designing new CUDA kernels, quantization-aware training and inference, and compilation/deployment techniques.
- Work with researchers and ML engineers on best-practices for optimal resource usage.
- Create and improve tooling and dashboards to ensure broad adoption of your work.
Qualifications:
- MS/PhD or Bachelors degree with a minimum of 4 years of industry experience in Computer Science, Robotics and/or similar technical field(s) of study.
- Solid coding proficiency in a variety of coding languages including Python, C++ or Rust.
- Experience in deep learning frameworks such as PyTorch or Jax.
- Skilled in profiling CPU and GPU code using tools such as PyTorch Profiler and NVIDIA Nsight.
- Open-minded and collaborative team player with willingness to help others.
- Passionate about self-driving technologies, solving hard problems, and creating innovative solutions.
Bonus/nice to have:
- Experience in identifying when custom CUDA kernels are needed, and implementing them.
- Experience in Bazel in a monorepo environment, and integrating third party packages into dev environments.
- Experience with Kubernetes-based training platforms.

