368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$79k – $90k per year
Location
Remote (Canada)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Smile Digital Health is a health information technology company headquartered in Toronto, Canada, and founded in 2016. The company provides a FHIR-native health data platform that enables interoperability, digital quality measures, and clinical intelligence for healthcare organizations. It serves government agencies, health systems, and payers globally, focusing on open standards to unify fragmented healthcare data and improve patient outcomes.

The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of production-grade services deployed across multiple cloud vendors and infrastructure platforms for Smile Digital Health, its clients, and partners.

This role designs and automates performance testing frameworks, integrates them into CI/CD pipelines, and uses observability tools to proactively detect and resolve bottlenecks. Working closely with engineering, product, and security teams, the SRE ensures systems meet strict SLAs for performance and availability while driving continuous optimization across multiple cloud platforms.

Responsibilities:

  • Collaborate with our Security Operations teams to define and implement best practices around Cloud Service Provider configuration for Azure and other cloud providers.
  • Develop, implement, and coordinate a multi-tenant approach around service offerings for databases, container platforms, authentication, certificates, and product registries.
  • Design, develop, and maintain cloud performance testing strategies, frameworks, and environments to validate application scalability, reliability, and resiliency.
  • Develop and automate load, stress, spike, and endurance (soak) testing as part of CI/CD pipelines.
  • Analyze application and infrastructure performance to identify bottlenecks and recommend performance optimizations across cloud-native services.
  • Develop and maintain cost and utilization tracking and attribution processes across Cloud Service Providers.
  • Create documentation detailing Cloud Service Provider offerings, implementation patterns, and best practices.
  • Develop and maintain technical relationships with our core Cloud Service Providers.
  • Implement and maintain secure, scalable infrastructure platforms for delivering cloud services.
  • Ensure internal and external SLAs are consistently met or exceeded, while continuously monitoring and improving system performance, reliability, and availability.
  • Create tools for automating deployment, monitoring, and platform operations.
  • Implement and manage observability solutions (logging, metrics, tracing) using OpenTelemetry, Prometheus, Grafana, Azure Monitor, and related technologies to provide actionable performance insights.
  • Plan and execute chaos engineering experiments to evaluate and improve application resiliency and fault tolerance.

Requirements:

    • 5+ years of experience with Cloud Service Providers and best practices around implementation and configuration, preferably managing Azure environments supporting SaaS products.
    • Experience working across multiple cloud providers (Azure required; AWS and/or Google Cloud Platform considered an asset).
    • Strong experience in Cloud Performance Engineering, including performance analysis, capacity planning, scalability testing, and optimization of distributed cloud-native applications.
    • Proven experience working with microservices architecture, with a strong focus on Java-based services.
    • Experience applying Chaos Engineering practices to evaluate and improve system resiliency.
    • Strong experience designing and executing performance testing strategies, including load, stress, spike, and endurance (soak) testing, to validate application scalability and defined latency and error-rate thresholds.
    • Hands-on experience with performance testing tools such as JMeter, Gatling, Azure Load Testing, or k6.
    • Experience validating application services sustaining 500+ transactions per second (TPS) while meeting defined performance objectives.
    • Hands-on experience deploying and managing containerized applications using Docker and Kubernetes, including autoscaling and performance optimization.
    • Experience using Terraform to provision and manage cloud infrastructure using Infrastructure as Code (IaC).
    • Experience tuning Kafka (partitioning, consumer group sizing, throughput/latency trade-offs) and other messaging/queueing platforms to sustain target transaction rates.
    • Hands-on experience implementing and using observability platforms including OpenTelemetry, Prometheus, Grafana, Azure Monitor, Application Insights, and Log Analytics.
    • Proven experience with Security and Compliance (SOC 2, HIPAA, ISO 27001) best practices and implementing controls that support high-velocity software delivery teams.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Toronto
$98k – $195k per year (Estimated) • In office • Full-Time • 7+ years exp • Wellington
Java
Python
SQL
Java
Spring Boot
Databases
Apache Kafka
Databricks
Neo4j
AI/ML
Flink
Spark
Frontend
GraphQL
DevOps
Azure
CI/CD
Datadog
Dynatrace
Kibana
Kubernetes
OpenShift
Platform Engineering
Splunk
Amazon ECS
Apply
$84k – $178k per year (Estimated) • In office • Full-Time • 10+ years exp • Wellington
Java
Python
DevOps
Ansible
AWS
Azure
CI/CD
Docker
GCP
Helm
Kubernetes
Platform Engineering
Prometheus
Service Mesh
Terraform
GitLab
IAM
Apply
Platform Engineer 1 day ago
$87k – $139k per year • In office • Full-Time • 3+ years exp • Berlin
Databases
PostgreSQL
Redis
DevOps
AWS
Azure
Bicep
CI/CD
Docker
GCP
GitHub Actions
Kubernetes
OpenShift
Terraform
GitHub
Apply
Founding Engineer 1 day ago
$81k – $116k per year • In office • Full-Time • Bachelor's Degree • Munich
JavaScript
Python
TypeScript
Databases
MySQL
PostgreSQL
Frontend
Next.js
React.js
Tailwind CSS
DevOps
AWS
Azure
CI/CD
Docker
GCP
Grafana
Kubernetes
OpenTelemetry
Prometheus
Apply
$175k – $195k per year • Remote • 8+ years exp
Python
Python
pySpark
AI/ML
Spark
Edge AI
DevOps
AWS
CI/CD
Incident Management
Platform Engineering
Amazon S3
Analytics
ETL/ELT
Apply
$61k – $79k per year • Remote • Full-Time • 8+ years exp • Toronto
Java
Java
Hibernate
DevOps
Azure
Azure DevOps
Git
Apply
$83k – $94k per year • Remote • Full-Time • Toronto
SQL
Databases
Apache Kafka
Azure SQL Database
DevOps
ArgoCD
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitOps
Grafana
Helm
Kubernetes
Loki
Nginx
Prometheus
Terraform
Traefik
Apply
$90k – $110k per year • Remote • 5+ years exp • Bachelor's Degree • Toronto
JavaScript
Node JS
SQL
TypeScript
Java
Java
Spring Framework
Frontend
Angular
npm
DevOps
CI/CD
Docker
Git
Apply
$65k – $79k per year • Remote • Full-Time • 5+ years exp • Toronto
JavaScript
Node JS
SQL
TypeScript
Java
Java
Spring Framework
Frontend
Angular
npm
DevOps
CI/CD
Docker
Git
Apply
$115k – $135k per year • Remote • 8+ years exp • Bachelor's Degree • Toronto
Java
SQL
DevOps
Git
Rest API
Apply
Actuarial Analyst 1 hour ago
$73k – $146k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Quebec • Waterloo • Toronto
Visual Basic
Apply
$47k – $109k per year (Estimated) • In office • Full-Time • Toronto
SQL
Visual Basic
Apply
$45k – $106k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Toronto
AI/ML
Copilot
Analytics
Power BI
Management
Power Apps
Apply
$48k – $112k per year (Estimated) • In office • Full-Time • Master's Degree • Toronto
C++
MATLAB
Python
Apply
Sr. UX Designer 1 hour ago
$76k – $161k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Toronto
Python
AI/ML
AI Agents
Hallucination
Human-in-the-Loop
LLM
LLM Guardrails
Design
Figma
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.