598,796open jobs
30,416companies
86,197added this week
Browse all
Salary
$90k – $180k per year
Location
In office (Toronto)
Seniority
Middle · 4+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match

About The Role

Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. We are looking for a Site Reliability Engineer to help build and operate the infrastructure behind that work.

Based in Toronto or remote, you will work across the systems that enable large-scale AI training and serving: high-performance networks, GPU clusters, storage, scheduling, and the operational tooling that keeps them reliable. This is a hands-on role for someone who enjoys taking complex infrastructure from “it works” to dependable, observable, and scalable.

You do not need to be an expert in every layer of the stack. We are looking for deep strength in at least one area-networking, cluster scheduling, storage, GPU systems, or AI infrastructure- and the curiosity and judgment to collaborate across the rest.

Responsibilities

  • Design, operate, and improve reliable infrastructure for AI training and inference workloads
  • Own and automate operational workflows across one or more core areas: networking, compute allocation, storage, GPU/server configuration, or AI platforms
  • Build monitoring, alerting, runbooks, and incident-response practices that make systems easier to operate
  • Diagnose performance, capacity, and reliability issues across hardware, operating systems, networks, schedulers, and distributed workloads
  • Partner closely with ML, research, and platform teams to translate workload needs into practical infrastructure improvements
  • Improve provisioning, configuration management, testing, and deployment automation
  • Help plan cluster growth, capacity allocation, upgrades, and lifecycle management
  • Contribute to a thoughtful reliability culture through documentation, post-incident learning, and pragmatic engineering standards

Minimum Qualifications

  • 4+ years of experience in site reliability engineering, infrastructure engineering, systems engineering, or a related production-operations role
  • Strong hands-on expertise in at least one of the following:
  • Networking, including firewalls, switching, routing, ASN/BGP configuration, or InfiniBand
  • Cluster and systems allocation with Kubernetes, SLURM, MAAS, or similar platforms
  • Distributed storage, particularly Ceph
  • GPU and server administration, including CUDA drivers, firmware, BIOS, and hardware troubleshooting
  • AI training or model-serving infrastructure
  • Experience operating production systems with a focus on availability, performance, security, and automation
  • Strong Linux administration and scripting skills
  • A systematic approach to troubleshooting across multiple layers of a complex system
  • Clear written and verbal communication skills, including the ability to work effectively with a distributed team

Preferred Qualifications

  • Experience supporting GPU-intensive AI or HPC environments

  • Experience with NVIDIA GPUs, CUDA, NCCL, and high-performance interconnects - Experience with InfiniBand, RDMA, RoCE, or 100Gb+ Ethernet

  • Familiarity with Kubernetes, SLURM, MAAS, Terraform, Ansible, or similar infrastructure tooling

  • Experience operating or tuning Ceph clusters

  • Familiarity with observability tooling such as Prometheus, Grafana, and centralized logging systems

  • Experience with hardware provisioning, firmware management, and bare-metal automation

  • Experience running large-scale distributed training or high-throughput inference workloads

  • Familiarity with cloud and hybrid infrastructure across AWS, GCP, or Azure

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
598,796 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Toronto
$29k – $72k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Bengaluru
Python
Bash
DevOps
gRPC
Terraform
Ansible
GCP
Jaeger
OpenTelemetry
Prometheus
Azure
CI/CD
GitOps
ArgoCD
AWS
Kubernetes
Grafana
Incident Management
GitLab
Apply
$35k – $81k per year (Estimated) • In office • Full-Time • 5+ years exp • Master's Degree • Mumbai
AI/ML
Fine-tuning
Transformers
TensorFlow
PyTorch
LLM
RAG
BERT
Hugging Face
DevOps
GCP
Azure
AWS
Docker
Kubernetes
Apply
$98k – $209k per year (Estimated) • In office • Full-Time • 3+ years exp • Palo Alto
Python
TypeScript
Bash
DevOps
Terraform
Ansible
GCP
GitHub Actions
CloudFormation
Pulumi
GitLab CI
Azure
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Apply
$26k – $61k per year (Estimated) • In office • 5+ years exp • Bengaluru
JavaScript
TypeScript
SQL
C#
C#
.NET
Databases
Azure Cosmos DB
Frontend
Redux
Webpack
Angular
React.js
Vite
Mobile
React Native
Clean Architecture
Dependency Injection
State Management
DevOps
Rest API
Azure DevOps
GitHub Actions
OpenTelemetry
Prometheus
Azure
CI/CD
Git
Docker
Kubernetes
Azure AKS
Management
Agile
Apply
.NET Developer 1 day ago
$26k – $61k per year (Estimated) • In office • 5+ years exp • Bengaluru
JavaScript
TypeScript
SQL
C#
C#
.NET
Databases
Azure Cosmos DB
Frontend
Redux
Webpack
React.js
Vite
Mobile
React Native
Clean Architecture
Dependency Injection
State Management
DevOps
Rest API
Azure DevOps
GitHub Actions
OpenTelemetry
Prometheus
Azure
CI/CD
Git
Docker
Kubernetes
Azure AKS
Management
Agile
Apply
Datacenter Technician 15 days ago
$36k – $72k per year • In office • Full-Time • Barrie
AI/ML
NLP
Apply
$150k – $270k per year • In office • Full-Time • Santa Clara
Python
Go
Java
Rust
C++
Databases
Apache Kafka
AI/ML
Spark
AI Agents
Flink
LLM
Edge AI
Agentic Workflows
DevOps
GCP
CI/CD
AWS
Docker
Kubernetes
Vector
Amazon Kinesis
Analytics
ETL/ELT
Apply
$180k – $400k per year • In office • Full-Time • Santa Clara
Python
Go
Java
Rust
C++
Databases
Apache Kafka
AI/ML
LangChain
Spark
LlamaIndex
Model Context Protocol
AI Agents
Flink
LLM
RAG
A2A
Edge AI
Agentic Workflows
DevOps
GCP
CI/CD
AWS
Kubernetes
Vector
Amazon Kinesis
Analytics
ETL/ELT
Apply
Frontend Engineer 1 year ago
$150k – $400k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • Santa Clara
JavaScript
TypeScript
Databases
Supabase
AI/ML
Multimodal AI
AI Agents
LLM
Edge AI
Frontend
Tailwind CSS
Next.js
D3.js
Chart.js
React.js
Apache ECharts
Mobile
Firebase
DevOps
WebRTC
WebSockets
CI/CD
GitHub
Management
Agile
Apply
$108k – $289k per year • In office • Full-Time • Bachelor's Degree • Toronto
Python
Rust
TypeScript
AI/ML
Fine-tuning
JAX
Multimodal AI
AI Agents
PyTorch
LLM
RAG
Edge AI
DevOps
GCP
WebRTC
Azure
AWS
GitHub
Apply
Deployment Engineer 5 hours ago
$40k – $116k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Toronto
Python
TypeScript
AI/ML
AI Agents
OpenAI
Scale AI
Apply
$24k per year • In office • Full-Time • High School Diploma • Toronto
Apply
$60k – $70k per year • Remote • Full-Time • 2+ years exp • Bachelor's Degree • Toronto
Apply
$52k – $100k per year (Estimated) • Remote/Hybrid • Contractor • 3+ years exp • Toronto
Analytics
A/B Testing
Microsoft Excel
Management
Slack
Jira
Apply
$32k – $68k per year (Estimated) • Remote/Hybrid • Part-Time • Toronto
AI/ML
OpenAI
Apply
See all jobs
This is one of many
598,796 more open roles from verified company boards, updated every day.