386,061open jobs
10,141companies
47,930added this week
Browse all
Salary
$95k – $205k per year (Estimated)
Location
Remote (Canada)
Seniority
Senior · 7+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is a Belgian recruitment platform built entirely around remote and flexible work, aggregating openings from thousands of employers that allow work from outside an office. Its matching engine ranks roles against a candidate's skills, seniority and stated preferences on location and flexibility, rather than leaving people to filter a keyword search, and it verifies how genuinely remote each posting is. The company also runs an AI screening layer that shortlists applicants for employers, and publishes research and guidance on distributed work practices alongside the job marketplace itself.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior DevOps Engineer, AI Platform based in Canada.

As a Senior DevOps Engineer, you’ll build and operate the cloud infrastructure powering AI platforms, web applications, APIs, and backend services.

You’ll translate technical designs into secure, scalable, reliable, and production-ready environments across modern cloud platforms.

The role combines deep Kubernetes expertise with cloud networking, infrastructure automation, CI/CD, and observability.

You’ll support AI workloads including LLM gateways, agent runtimes, RAG pipelines, and asynchronous processing services.

You’ll work closely with AI engineers, application developers, and architects while independently owning infrastructure delivery and operations.

Production reliability, incident response, scalability, security, and cost optimization will be central to your impact.

This is an opportunity to help create reusable platform capabilities that enable engineering teams to deliver sophisticated services faster and more consistently.

Accountabilities

    • Translate application and platform technical designs into reliable, secure, scalable, and production-ready cloud infrastructure with minimal supervision.
    • Design, provision, operate, and troubleshoot Kubernetes environments, primarily using Azure Kubernetes Service and Oracle Kubernetes Engine.
    • Build and support infrastructure for AI workloads, including LLM gateways, Python-based agent runtimes, RAG workers, MCP services, background workers, and asynchronous processing pipelines.
    • Design and manage cloud networking, including virtual networks, subnets, routing, NAT, load balancers, DNS, TLS, private connectivity, firewalls, network policies, ingress, egress, and service-to-service communication.
    • Operate infrastructure supporting web applications, APIs, databases, caches, queues, scheduled jobs, microservices, and event-driven workloads.
    • Build and maintain CI/CD pipelines using Jenkins and Bitbucket, integrating Docker, Helm, Kubernetes, ArgoCD, and container registries.
    • Automate infrastructure provisioning and configuration using Terraform, Helm, Kubernetes manifests, Python, Bash, and Infrastructure as Code practices.
    • Implement comprehensive observability across infrastructure and applications through metrics, logs, distributed tracing, dashboards, alerts, health checks, and service-level objectives.
    • Own production readiness, incident troubleshooting, root cause analysis, scalability, reliability, and infrastructure cost optimization.
    • Create reusable infrastructure patterns, templates, and operational practices that enable engineering teams to launch services efficiently and consistently.
    • Determine required cloud resources, Kubernetes configurations, namespaces, scaling models, identities, secrets, and supporting services for new workloads.
    • Provision and operate dependencies such as PostgreSQL, Redis, RabbitMQ, storage systems, and other shared platform services.
    • Establish CI/CD workflows covering builds, testing, container publishing, deployment, validation, and rollback.
    • Define operational runbooks, capacity monitoring, dashboards, alerts, and health checks before production launches.
    • Own infrastructure delivery through UAT and production while partnering with architects and engineers to resolve technical design trade-offs.
    • Requirements

      • 7+ years of professional experience in DevOps, Site Reliability Engineering, Platform Engineering, Cloud Infrastructure, or a closely related discipline.
      • Strong hands-on experience operating production Kubernetes environments, including expertise in networking, scheduling, storage, autoscaling, security, and troubleshooting.
      • Strong Microsoft Azure experience, particularly with AKS, networking, identity, storage, and monitoring; Oracle Cloud Infrastructure experience is preferred.
      • Deep understanding of cloud networking concepts, including virtual networks, subnets, routing, NAT, load balancing, private networking, DNS, TLS, firewalls, ingress, and egress.
      • Proven experience with Jenkins, Bitbucket, Docker, Terraform, Helm, Kubernetes, and Infrastructure as Code.
      • Experience supporting production web applications and backend services, including REST APIs, microservices, background workers, and asynchronous architectures.
      • Hands-on knowledge of databases, caching, and messaging technologies such as PostgreSQL, Redis, RabbitMQ, or equivalent platforms.
      • Experience implementing production observability using tools such as OpenTelemetry, Grafana, Prometheus, Sentry, or cloud-native monitoring solutions.
      • Strong Linux, systems administration, and production troubleshooting capabilities.
      • Working knowledge of Python, particularly backend services built with frameworks such as FastAPI.
      • Familiarity with at least one additional programming language such as C#, Java, Go, JavaScript, or TypeScript.
      • Strong understanding of HTTP/HTTPS, DNS, TCP/IP, proxies, authentication, APIs, connection pooling, caching, concurrency, queues, retries, dead-letter queues, and asynchronous processing.
      • Ability to read application logs and stack traces and diagnose infrastructure and application issues involving latency, memory, CPU, connections, and dependencies.
      • Strong cross-functional communication skills and the ability to independently execute technical designs while engaging architects and application engineers when needed.
      • Experience supporting AI or machine learning platforms, LLM gateways, agent runtimes, RAG pipelines, or MCP services is highly desirable.
      • Familiarity with Cloudflare, Envoy, ArgoCD, GitOps, and OpenTelemetry would be an advantage.
      • Experience building reusable infrastructure platforms for high-scale SaaS or customer-facing applications is a plus.
      • Strong experience operating distributed systems using technologies such as RabbitMQ, Redis, and PostgreSQL is beneficial.
      • Benefits

        • Full-time, fully remote position available across Canada.
        • Opportunity to work on infrastructure supporting modern AI platforms, agent systems, RAG workloads, and cloud-native applications.
        • High-impact role with significant ownership over production infrastructure, reliability, scalability, and platform engineering.
        • Opportunity to work across Microsoft Azure and Oracle Cloud Infrastructure.
        • Exposure to modern technologies including Kubernetes, Terraform, Helm, ArgoCD, Docker, OpenTelemetry, and GitOps.
        • Collaborative environment working closely with AI engineers, application engineers, architects, and platform teams.
        • Opportunity to create reusable infrastructure capabilities that accelerate engineering delivery.
        • Professional growth through hands-on work with large-scale distributed systems and emerging AI infrastructure.
        • Flexible remote-first working environment designed to support collaboration across distributed teams.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
386,061 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$36k – $44k per year • Remote • Full-Time • 2+ years exp • High School Diploma
Bash
PHP
PHP
WordPress
Apply
$47k – $54k per year • Equity • Remote • Full-Time • 3+ years exp
C#
SQL
Databases
MS SQL
DevOps
Azure
Azure DevOps
GitHub
Management
Jira
Apply
$47k – $54k per year • Equity • Remote • Full-Time • 3+ years exp
C#
SQL
Databases
MS SQL
DevOps
Azure
Azure DevOps
GitHub
Management
Jira
Apply
$79k – $168k per year (Estimated) • Remote • Full-Time • 6+ years exp • Bachelor's Degree
Go
Java
Python
Databases
PostgreSQL
AI/ML
Flink
LLM
Spark
DevOps
AWS
Azure
Bazel
CI/CD
Docker
Envoy
GCP
GitOps
Grafana
gRPC
Istio
Kubernetes
Linkerd
OpenTelemetry
Platform Engineering
Prometheus
Service Mesh
SRE
Terraform
Cybersecurity
GDPR
Keycloak
SOC 2
Apply
$111k – $215k per year (Estimated) • Remote • Full-Time • 6+ years exp • Bachelor's Degree
Go
Java
Python
Databases
PostgreSQL
AI/ML
Flink
LLM
Spark
DevOps
AWS
Azure
Bazel
CI/CD
Docker
Envoy
GCP
GitOps
Grafana
gRPC
Istio
Kubernetes
Linkerd
OpenTelemetry
Platform Engineering
Prometheus
Service Mesh
SRE
Terraform
Cybersecurity
GDPR
Keycloak
SOC 2
Apply
$94k – $168k per year (Estimated) • Remote • Full-Time • 7+ years exp
Apply
$130k – $222k per year (Estimated) • Remote • Full-Time • 7+ years exp
Apply
$36k – $44k per year • Remote • Full-Time • 2+ years exp • High School Diploma
Bash
PHP
PHP
WordPress
Apply
$115k – $129k per year • Remote • Full-Time • 10+ years exp
PowerShell
Python
SQL
C#
Java
C#
.NET
Entity Framework Core
Java
Flyway
Liquibase
Databases
Amazon Aurora
MS SQL
Oracle
PostgreSQL
AI/ML
AI Agents
DevOps
Amazon CloudWatch
Amazon EC2
AWS
CloudFormation
IAM
SLI/SLO/SLA
Terraform
Cybersecurity
SOC 2
Apply
$47k – $54k per year • Equity • Remote • Full-Time • 5+ years exp
SQL
Databases
MS SQL
Apply
See all jobs
This is one of many
386,061 more open roles from verified company boards, updated every day.