Salary
≈ $162k – $332k per year (Estimated)
Location
Hybrid (United States, Canada)
Seniority
Staff · 5+ years exp
Employment
Full-Time
Confirmed on the employer's own hiring board on Sep 24, 2026. First seen by Alion on Sep 23, 2026. Waabi scores A on the Alion truth index.
Overview
Company
Impact
Profile match
Waabi is an artificial intelligence and autonomous vehicle technology company headquartered in Toronto, Canada, founded in 2021. The company develops an end-to-end artificial intelligence software platform, known as the Waabi Driver, and a simulation environment called Waabi World to power autonomous trucks and robotaxis. It operates primarily in the North American logistics and transportation sectors through strategic partnerships with industry leaders such as Volvo, Uber Freight, and NVIDIA.
You will..
- Build and evolve our training infrastructure on Kubernetes with Infrastructure - GPU scheduling, autoscaling, multi-node distributed jobs, capacity strategy, and the operators and workflow engines that keep long-running training reliable.
- Shape the developer-facing surface - CLIs, SDKs, job submission, templates, paved paths - designed with the teams who'll use them. Make the common case one command and keep the uncommon case possible.
- Shorten the inner loop. Time to first training run, edit-to-signal latency, local iteration before a job hits the cluster, fast failure over slow mystery. Measure it, publish it, drive it down.
- Evangelize best-in-class tooling and frameworks. Track what the ecosystem is shipping, evaluate honestly, and make the case with working prototypes and migration paths - or say plainly when a shiny thing isn't worth the switching cost.
- Strengthen the data and artifact layer. Dataset versioning, sharding, and high-throughput loading of large multimodal sensor data, so jobs saturate GPUs instead of waiting on I/O.
- Turn one-off Python into durable tooling - tested, documented, observable libraries, CLIs, and services with sane defaults, and deletions where they're overdue.
- Make experiments legible, with the teams who live in them: experiment hygiene, dashboards researchers trust, a real model registry, and lineage from dataset to checkpoint to simulation result.
- Ship CI/CD for models alongside autonomy and simulation, so a model change is validated the same way a code change is.
- Build observability across the ML stack - utilization, throughput, failure modes, queue times, cost per experiment. When a job fails at 3am on node 47, the researcher should find out why without you.
- Treat docs, onboarding, and support as product surface - golden-path guides, a new researcher productive on day two, office hours that turn repeat questions into shipped fixes.
- Drive adoption, not just availability. Prototype with real users, watch them work, iterate. A tool nobody adopts didn't ship.
- Make the platform boringly reliable - fewer failures, faster recovery, and none of the manual steps that quietly cost a team days.
- Build guardrails that don't feel like walls, with Security, IT, and Infrastructure: access controls, data handling, and cost governance that hold up in an IP-sensitive environment while staying self-serve.
Qualifications:
- 5+ years of software or infrastructure engineering, including tools or platforms used by other engineers and operating ML or data-intensive production systems.
- Hands-on Kubernetes expertise - GPU scheduling, autoscaling, Helm or equivalent, networking fundamentals, and the ability to debug a cluster under load rather than restart it.
- Excellent Python, and a track record of designing APIs and CLIs other people enjoy using.
Practical AWS depth: object storage at scale, IAM, GPU compute, networking, cost management, and infrastructure as code (Terraform, Pulumi, or similar).
- Distributed training in PyTorch (DDP, FSDP, or similar), plus experiment tracking and model registry tooling - from the perspective of someone who made them pleasant for others to use.
- Fluency with containers, CI/CD, and modern build systems, including large monorepos.
- The ability to influence without authority: evaluate a framework on its merits, pilot it credibly, and persuade skeptical senior engineers to change how they work.
- A collaborative default - you'd rather co-own a system than draw a boundary around your part of it.
- User empathy: you'd rather fix the third-most-interesting problem blocking ten people than the most interesting one blocking nobody.
- Strong product instincts, strong writing, and comfort operating autonomously in ambiguous territory.
- Passionate about self-driving technologies and frontier AI, and about what a small, world-class team can do with the right infrastructure.
Bonus/nice to have:
- Internal developer platform, research platform, or DevEx work - with a story about a tool whose adoption you grew from zero.
- Large-scale distributed GPU training: hundreds to thousands of accelerators, NCCL, high-performance cluster networking, collective communication tuning.
- High-throughput loading of LiDAR or camera data, and formats such as Parquet or WebDataset.
- Workflow and scheduling systems - Argo Workflows, Ray, Flyte, Kubeflow, or Slurm.
- Build-system depth (Bazel or similar), including remote caching in a monorepo.
- Simulation infrastructure or large-scale batch evaluation pipelines.
- Background in ML, robotics, or autonomous systems infrastructure.
- Security- and IP-sensitive production environments.
- Open-source contributions to ML infrastructure or developer tools.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
745,032 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Free forever. No card. Under a minute.
Your match
How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.
Recommended for you based on this role
AI/ML
Similar stack
Same company
In your city
CXO AI Engineer
2 hours ago
≈ $152k – $265k per year (Estimated) • Equity • Remote (United States) • 5+ years exp
Python
SQL
Databases
Amazon Redshift
Trino
AI/ML
Model Context Protocol
dbt
Function Calling
AI Agents
LLM
RAG
Context Engineering
Tool Use
DevOps
SLI/SLO/SLA
Management
Slack
Apply
Apply
Principal AI Architect
1 hour ago
$270k – $330k per year • In office • Full-Time • 12+ years exp • New York
AI/ML
Machine Learning
Apply
$133k – $302k per year • Hybrid • Full-Time • 12+ years exp • Associate's Degree • New York • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
AWS Bedrock
AWS Bedrock AgentCore
DevOps
AWS
Apply
Agentic Operations Engineer 6491979
1 hour ago
$74k – $220k per year • Remote (United States) • Full-Time • 7+ years exp • High School Diploma • Philadelphia
Python
Java
AI/ML
Copilot
Cursor
LangGraph
AutoGen
LangChain
Claude Code
Prompt Engineering
AI Agents
Semantic Kernel
CrewAI
RAG
Multi-Agent Systems
DevOps
Rest API
GCP
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Apply
Senior Full Stack Engineer
2 hours ago
≈ $88k – $176k per year (Estimated) • Hybrid • Full-Time • London
Python
JavaScript
Python
Django
Django REST Framework
Databases
PostgreSQL
OpenSearch
Frontend
React.js
DevOps
Terraform
GCP
Azure
AWS
Shift-Left
Cybersecurity
Shift-Left Security
Apply
Quantitative Engineer
2 hours ago
$90k – $156k per year • In office • Full-Time • 1+ year exp • Bachelor's Degree • Chicago
Python
JavaScript
TypeScript
Python
pySpark
AI/ML
Hadoop
Spark
Pandas
Machine Learning
Frontend
Angular
React.js
Apply
Sr Accountant (Revenue & AR Reserves)
2 hours ago
≈ $69k – $134k per year (Estimated) • Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • Raleigh
Python
SQL
Databases
Snowflake
Databricks
DevOps
Azure
Analytics
Alteryx
Microsoft Excel
Management
Power Automate
Apply
Electronics Technician
2 hours ago
≈ $44k – $81k per year (Estimated) • Hybrid • Full-Time • 4+ years exp • Canada
Python
DevOps
Linux
Windows
Management
Outlook
Microsoft Office
Apply
Senior Software Engineer, Fullstack
2 hours ago
≈ $130k – $241k per year (Estimated) • Hybrid • Full-Time • 1+ year exp • United States
JavaScript
TypeScript
Node JS
Frontend
GraphQL
Next.js
React.js
Sass
DevOps
Rest API
Terraform
Azure
AWS
TCP/IP
QA
Playwright
Jest
Vitest
Apply
≈ $166k – $361k per year (Estimated) • Hybrid • Full-Time • 4+ years exp • Bachelor's Degree
Python
C++
C++
PyTorch C++
AI/ML
PyTorch
Machine Learning
Apply
Senior / Staff ML Training Optimization Engineer
4 months ago
≈ $150k – $335k per year (Estimated) • Hybrid • Full-Time • 4+ years exp • Bachelor's Degree
Python
Rust
C++
C++
PyTorch C++
AI/ML
CUDA Toolkit
Quantization
PyTorch
CUDA
DevOps
Kubernetes
Bazel
Apply
Research Scientist, Simulation Agents
6 months ago
≈ $108k – $289k per year (Estimated) • Hybrid • Full-Time • Master's Degree
AI/ML
Reinforcement Learning
TensorFlow
PyTorch
Machine Learning
Robotics
Imitation Learning
Reinforcement Learning
Apply
≈ $150k – $335k per year (Estimated) • Hybrid • Full-Time • 6+ years exp • Bachelor's Degree
Python
Rust
C++
C++
PyTorch C++
AI/ML
CUDA Toolkit
TensorRT
PyTorch
CUDA
Apply
≈ $103k – $276k per year (Estimated) • Hybrid • Full-Time • Bachelor's Degree
Python
Rust
C++
C++
PyTorch C++
AI/ML
CUDA Toolkit
Fine-tuning
Reinforcement Learning
Computer Vision
TensorRT
PyTorch
CUDA
Machine Learning
Robotics
Motion Planning
Reinforcement Learning
Apply
This is one of many
745,032 more open roles from verified company boards, updated every day.

