368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$120k – $244k per year (Estimated)
Location
Remote (Canada)
Seniority
Staff · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Caseware is a global leader in providing software solutions for accounting, auditing, and financial reporting. Their products, such as Caseware Cloud and Caseware Working Papers, are designed to modernize accounting practices, streamline audits, and provide robust financial reporting tools for accounting firms, corporations, and government entities. Caseware is renowned for integrating advanced technologies, including AI-driven tools, to enhance efficiency and insight for its users.

This is a hands-on senior engineering role focused on improving production resilience, strengthening security, driving operational excellence, and enhancing the developer experience across the organization.

In this role, you will design, build, and evolve the foundational systems, tooling, and operational practices that enable engineering teams to ship secure, reliable, and scalable software with confidence. You will help establish reliability standards, define service level objectives (SLOs), improve observability, automate operational processes, and drive incident management and post-incident learning practices that strengthen platform stability over time.

Partnering closely with Engineering, Security, Platform, and Product teams, you will architect scalable distributed systems, optimize Kubernetes and AWS-based infrastructure, and build automated delivery pipelines that support rapid and safe software releases. You will play a key role in reducing operational toil, improving system performance, increasing platform reliability, and ensuring that our infrastructure can support continued business growth.

This is a full-time permanent position

This is an existing vacancy

Location:This is a remote location open to candidates legally authorized to work in Canada.

What you will be doing:

  • Drive reliability engineering initiatives and operational excellence for mission-critical services running on AWS and Kubernetes.
  • Design, implement, and continuously improve deployment, release, and rollback strategies across complex distributed systems.
  • Establish secure-by-default CI/CD pipelines with robust automation, governance, and policy-driven controls.
  • Enhance platform observability through metrics, logs, tracing, and actionable alerting to improve system visibility and operational efficiency.
  • Define, implement, and mature Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability standards across the organization.
  • Lead response efforts for high-severity incidents, ensuring timely resolution, effective communication, and meaningful post-incident reviews that drive continuous improvement.
  • Partner closely with engineering teams to strengthen platform standards, improve service resilience, optimize runtime performance, and embed reliability best practices.
  • Mentor and guide engineers on cloud-native technologies, site reliability engineering principles, and operational excellence practices, fostering a culture of continuous learning and accountability.

What you will bring:

  • 8+ years of experience in Site Reliability Engineering (SRE), Platform Engineering, DevOps, or related cloud-native engineering roles.
  • Deep expertise in AWS services, including EKS, IAM, VPC, Lambda, CloudFront, S3, and cloud networking/security best practices.
  • Advanced experience operating and scaling production Kubernetes environments.
  • Strong hands-on experience with Istio service mesh, including traffic management, security, observability, and resiliency.
  • Proven expertise with Infrastructure as Code (IaC), preferably using AWS CDK.
  • Experience building and managing CI/CD pipelines using GitHub Actions or similar platforms.
  • Strong troubleshooting, performance optimization, and incident management experience in distributed systems.
  • Excellent communication, collaboration, and technical leadership skills.
  • Observability & Reliability

  • Experience designing and operating monitoring, logging, tracing, and alerting solutions for cloud-native platforms.
  • Strong knowledge of AWS CloudWatch, OpenTelemetry, AWS X-Ray, and Kubernetes observability tooling.
  • Experience defining and operationalizing SLIs, SLOs, alerting strategies, runbooks, and reliability metrics.
  • Proven ability to leverage observability data to improve service reliability, reduce incident impact, and optimize operational performance.
  • Software Engineering & Platform Development

  • Strong proficiency in TypeScript and Node.js for platform engineering, automation, and operational tooling.
  • Experience building and maintaining scalable backend services, APIs, and event-driven systems.
  • Deep understanding of Kubernetes architecture, controllers, Gateway API, ingress management, and service networking.
  • Experience implementing zero-trust architectures, mTLS, and service-to-service security controls.
  • Commitment to high-quality engineering practices, including automated testing, code reviews, and observability-driven development.
  • Strong understanding of resilience engineering, including autoscaling, disruption management, failure testing, and safe deployment strategies.
  • Nice to Have

  • Experience with progressive delivery practices such as canary, blue/green, and feature-flag-based deployments.
  • Experience working in regulated, compliance-driven, or security-sensitive SaaS environments.
  • Familiarity with FinOps principles and cost optimization strategies for cloud platforms.
  • Experience building internal developer platforms and self-service engineering tooling.
  • Cloud-native certifications such as CKA, CKAD, CKS, KCSA, or KCNA.
  • Kubestronaut certification or equivalent advanced Kubernetes expertise is highly regarded.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Toronto
$25k – $42k per year • Equity 0–0.2% • Remote • Full-Time • 3+ years exp
Bash
Go
JavaScript
Python
TypeScript
DevOps
AWS
Azure
CI/CD
Datadog
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Incident Management
Kubernetes
Platform Engineering
Prometheus
Terraform
Amazon CloudWatch
GitHub
GitLab
IAM
Cybersecurity
Least Privilege
Apply
$100k – $210k per year • Equity 0–0.5% • Remote • Full-Time • 3+ years exp • San Francisco
Bash
Go
JavaScript
Python
TypeScript
DevOps
AWS
Azure
CI/CD
Datadog
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Incident Management
Kubernetes
Platform Engineering
Prometheus
Terraform
Amazon CloudWatch
GitHub
GitLab
IAM
Cybersecurity
Least Privilege
Apply
$49k – $173k per year (Estimated) • In office • 2+ years exp • Bachelor's Degree • Kfar Saba
Java
Kotlin
Python
SQL
C#
C#
.NET
DevOps
AWS
CI/CD
Apply
Team Lead DevOps 1 day ago
$23k – $62k per year (Estimated) • Remote • 5+ years exp • Moscow
Bash
Python
Erlang
Erlang
EMQX
Databases
Apache Kafka
ClickHouse
PostgreSQL
RabbitMQ
Redis
Redpanda
Trino
DevOps
Ansible
AWS
AWX
FinOps
HAProxy
Hetzner
Kubernetes
SLI/SLO/SLA
Terraform
Yandex Cloud
Amazon S3
Apply
$100k – $200k per year • Equity 0.5–5% • In office • Full-Time • 1+ year exp • New York
Python
TypeScript
JavaScript
Python
FastAPI
Databases
DynamoDB
PostgreSQL
AI/ML
Claude
LLM
OpenAI
AI Agents
Frontend
Next.js
Tailwind CSS
React.js
DevOps
AWS
Docker
Vercel
GitHub
Management
Slack
Apply
$114k – $227k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • Toronto
Apply
$109k – $233k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Toronto
Go
AI/ML
AI Agents
Model Context Protocol
Apply
$71k – $202k per year (Estimated) • Remote • Full-Time • 8+ years exp • PhD
Python
TypeScript
Databases
DynamoDB
AI/ML
Hybrid Search
LLM
RAG
AI Agents
Human-in-the-Loop
ISO 42001
LLM Evaluation
LLM Guardrails
LLMOps
NIST AI RMF
Frontend
GraphQL
DevOps
Amazon EKS
AWS
AWS Lambda
CI/CD
CloudFormation
Docker
GitHub Actions
Incident Management
Kubernetes
New Relic
OpenTelemetry
Prometheus
Terraform
Vector
Amazon CloudWatch
Amazon S3
GitHub
IAM
Management
Confluence
Jira
Slack
Apply
$71k – $202k per year (Estimated) • Remote • Full-Time • 8+ years exp • PhD • Bogotá
Python
TypeScript
Databases
DynamoDB
AI/ML
Hybrid Search
LLM
RAG
AI Agents
Human-in-the-Loop
ISO 42001
LLM Evaluation
LLM Guardrails
LLMOps
NIST AI RMF
Frontend
GraphQL
DevOps
Amazon EKS
AWS
AWS Lambda
CI/CD
CloudFormation
Docker
GitHub Actions
Incident Management
Kubernetes
New Relic
OpenTelemetry
Prometheus
Terraform
Vector
Amazon CloudWatch
Amazon S3
GitHub
IAM
Management
Confluence
Jira
Slack
Apply
SDET II 4 days ago
$38k – $119k per year (Estimated) • Remote • Full-Time • 2+ years exp
JavaScript
TypeScript
AI/ML
Claude
Copilot
AI Agents
Devin
DevOps
AWS
CI/CD
GitHub
QA
Cypress
k6
Pact
Apply
Actuarial Analyst 26 min ago
$73k – $146k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Quebec • Waterloo • Toronto
Visual Basic
Apply
$47k – $109k per year (Estimated) • In office • Full-Time • Toronto
SQL
Visual Basic
Apply
$45k – $106k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Toronto
AI/ML
Copilot
Analytics
Power BI
Management
Power Apps
Apply
$48k – $112k per year (Estimated) • In office • Full-Time • Master's Degree • Toronto
C++
MATLAB
Python
Apply
Sr. UX Designer 1 hour ago
$76k – $161k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Toronto
Python
AI/ML
AI Agents
Hallucination
Human-in-the-Loop
LLM
LLM Guardrails
Design
Figma
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.