976,915open jobs
58,722companies
161,793added this week
Browse all
Salary
≈ $141k – $275k per year (Estimated)
Location
Remote (Canada)
Seniority
Senior · 10+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 30, 2026. First seen by Alion on Sep 28, 2026. NVIDIA scores A on the Alion truth index.

Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

We are looking for a Senior Software Engineer to help build NeMo Platform, NVIDIA’s product for developing, evaluating, deploying, and operating AI systems at scale. This role is for a senior engineer/architect for our Core team which owns and ships an open source plugin-based AI platform for running and optimizing Agents targeting multiple compute backends (local/docker, Kubernetes, Slurm, etc.).

As AI systems become more autonomous and more deeply integrated into real workflows, teams need robust APIs and orchestration systems for running, monitoring, and optimizing Agents at scale. Increasingly, the users of these systems are themselves autonomous or semi-autonomous Agents. The NeMo Platform group is building a sophisticated agent execution framework to enable agents to automatically run hundreds of experiments in parallel to find the most efficient agent architectures for our customers. This is important product engineering research for making agents more sustainable. AI systems are not yet nearly as efficient as they can be, and systems like NeMo Platform will allow large scale AI consumers to automatically tune their agents to use fewer tokens and rely on more efficient models with better throughput.

What you'll be doing:

  • Working in a product research environment where we place big bets on where the future is heading, adapting in real time as we build alongside an industry that is constantly evolving with us. This means fast iteration, high ownership, pragmatic decisions, and performance-minded implementation under production constraints

  • Designing an Agentic Execution system that are flexible enough to work in many environments (local, k8s, Slurm, on prem / air-gapped)

  • Provide senior technical leadership through design reviews, code reviews, mentoring, and ownership of ambiguous cross-component problems

  • Building and maintain our Core Platform APIs for running jobs, storing data, entities, secrets, and RBAC and Auth

  • Extending our flexible Plugin Architecture that makes it easy for many teams and external customers to install new capabilities into our system

  • Building in the open in our OSS repo, keeping up the high standards that the open source community demands

  • Shipping code at the speed of light with an unlimited token budget using best in class agentic coding tools

  • Improving reliability, observability, debuggability, and performance across NeMo Platform, SDKs, plugins, jobs, and developer workflows

  • Building strong test coverage across unit, integration, E2E, Docker, and Kubernetes workflows

What we need to see:

  • BS, MS, or equivalent experience in Computer Science, Computer Engineering, or a related technical field

  • 10+ years of professional software engineering experience building production systems

  • Comfort working in a very fast and ambiguous environment

  • Exceptional communication, both verbal and written. This includes the ability to produce and review high quality architectural RFCs, and to discuss them clearly with the right level of technical detail for the right people (Engineer, Product, Marketing, etc.)

  • Strong system design skills, with a pragmatic flexibility and phenomenal instincts to invent robust systems quickly without over-complicating. Strong understanding of reliability, scalability, security, and performance tradeoffs in production infrastructure

  • Experience with distributed systems, cloud-native services, containers, Kubernetes, and job orchestration

  • Excellent Python engineering skills, including API design, typing, testing, debugging, performance analysis, and maintainable software design

  • Experience designing SDKs, libraries, plugins, CLIs, or other developer-facing interfaces

  • Ability to work independently, define technical scope, break down ambiguous problems, and drive work across team boundaries

Ways to stand out from the crowd:

  • Experience building, deploying, and iterating on production agentic AI systems at scale in Kubernetes

  • Experience with sophisticated plugin architectures

  • Strong ability to connect technical evaluation work to business outcomes, product quality, user experience, reliability, or operational efficiency

  • Experience with enterprise AI systems where measurement, regression testing, observability, governance, and continuous improvement are required for production deployment

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 170,000 CAD - 220,000 CAD for Level 4, and 225,000 CAD - 275,000 CAD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until October 2, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
976,915 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
Toronto
≈ $88k – $179k per year (Estimated) • In office • Full-Time • 7+ years exp • Toronto
JavaScript
Java
TypeScript
SQL
Java
Spring Boot
Hibernate
Spring Cloud
Databases
MySQL
PostgreSQL
Apache Kafka
AI/ML
Claude Code
AI Agents
Frontend
GraphQL
React.js
DevOps
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Apache HTTP Server
Apply
$160k – $170k per year • Remote (United States) • Full-Time • Atlanta
Python
JavaScript
SQL
Python
Django
Django REST Framework
Databases
PostgreSQL
Redis
DevOps
Rest API
GitHub Actions
CI/CD
Git
AWS
Docker
GitHub
Amazon S3
Management
Stripe
QA
Playwright
Sentry
Apply
≈ $175k – $292k per year (Estimated) • Remote (United States) • 8+ years exp • Bachelor's Degree • Cambridge
Python
Go
Java
C++
Databases
Google Cloud Spanner
DevOps
Terraform
CI/CD
Kubernetes
Apply
$180k – $240k per year • Remote (United States) • Top Secret • 5+ years exp • Bachelor's Degree • Washington
Python
JavaScript
AI/ML
Computer Vision
Machine Learning
DevOps
GCP
Azure
AWS
Docker
Kubernetes
Cybersecurity
SOC 2
FedRAMP
Apply
$160k – $200k per year • Remote (United States) • 7+ years exp • Overland Park
JavaScript
Swift
Objective-C
AI/ML
AI Agents
Agentic Workflows
Frontend
React.js
Mobile
React Native
Offline-First
DevOps
AWS
Apply
In office
Python
SQL
PowerShell
Groovy
Databases
MySQL
PostgreSQL
DevOps
Rest API
Splunk
Terraform
Puppet
Ansible
GCP
OpenShift
Azure DevOps
GitHub Actions
Podman
Chef
CloudFormation
Dynatrace
Kong
Prometheus
GitLab CI
Azure
CI/CD
GitOps
Windows Server
Jenkins
Git
AWS
Docker
Kubernetes
Grafana
Blue-Green Deployment
Configuration Management
AppDynamics
Bicep
Bitbucket
JFrog Artifactory
Amazon EKS
Google GKE
Azure AKS
Incident Management
SLI/SLO/SLA
GitHub
IAM
API Gateway
Windows
DNS
VPN
Cybersecurity
Snyk
SonarQube
Trivy
Checkmarx
PCI DSS
Fortify
Management
Confluence
Jira
Agile
Scrum
Kanban
Apply
In office
JavaScript
TypeScript
SQL
Node JS
Databases
MySQL
PostgreSQL
Frontend
React.js
DevOps
Terraform
GCP
Azure
CI/CD
Docker
Kubernetes
Incident Management
Management
Agile
Apply
In office
Python
AI/ML
MLFlow
Kubeflow
Metaflow
Machine Learning
DevOps
Terraform
Ansible
GitHub Actions
CircleCI
CloudFormation
Fluentd
Prometheus
GitLab CI
CI/CD
Jenkins
Git
AWS
Grafana
Bitbucket
GitHub
GitLab
Apply
In office
DevOps
Azure DevOps
Azure
CI/CD
Kubernetes
Apply
Hybrid
SQL
DevOps
GCP
Azure
CI/CD
AWS
Docker
Linux
Unix
DNS
Apply
≈ $86k – $133k per year (Estimated) • Remote (Poland) • Full-Time • 6+ years exp • Bachelor's Degree • Poland • Czech Republic • Romania • Hungary
Rust
C++
AI/ML
CUDA Toolkit
CUDA
DevOps
Linux
TCP/IP
Apply
$224k – $357k per year • Remote (United States) • Full-Time • 12+ years exp • Bachelor's Degree • United States
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
CUDA Toolkit
TensorFlow
PyTorch
Synthetic Data
CUDA
Machine Learning
Apply
≈ $29k – $83k per year (Estimated) • Remote (India) • Full-Time • 3+ years exp • Bachelor's Degree • Bengaluru
Python
Java
Rust
Bash
AI/ML
NVIDIA NIM
DevOps
gRPC
CI/CD
ArgoCD
Kubernetes
GitLab
Linux
Unix
Apply
≈ $107k – $294k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Yokneam • Tel Aviv
Python
Java
DevOps
Kubernetes
Linux
Apply
≈ $110k – $301k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Tel Aviv
Python
C++
AI/ML
InfiniBand
DevOps
HPC
Linux
TCP/IP
Apply
≈ $85k – $168k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Toronto
Python
Java
SQL
COBOL
Java
Spring Boot
Hibernate
COBOL
IBM MQ
Databases
Apache Kafka
AI/ML
Polars
Pandas
Mobile
Dependency Injection
DevOps
Rest API
GitHub Actions
Jenkins
Docker
Kubernetes
Shift-Left
GitHub
Cybersecurity
Shift-Left Security
Management
Confluence
Jira
Agile
Scrum
Microsoft Office
Apply
≈ $78k – $159k per year (Estimated) • In office • Full-Time • Toronto
JavaScript
Java
Java
Spring Boot
DevOps
CI/CD
Kubernetes
Management
Agile
Apply
In office • Full-Time • Toronto
AI/ML
AI Agents
Apply
≈ $57k – $108k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Toronto
Python
SQL
Python
pySpark
AI/ML
Spark
DevOps
Azure
Management
Agile
Microsoft Office
Apply
$55k – $90k per year • In office • Full-Time • 1+ year exp • Toronto
Apply
See all jobs
This is one of many
976,915 more open roles from verified company boards, updated every day.