368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$88k – $183k per year (Estimated)
Location
Remote/Hybrid (Canada)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is an AI-powered job platform focused on remote and flexible work. It matches candidates with relevant roles using skills and preference-based algorithms, and also offers career coaching and job-search guidance.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in Canada.

This is an impactful opportunity for an experienced SRE to strengthen the reliability, scalability, security, and operational efficiency of modern cloud platforms.

You will take ownership of production systems across AWS and Kubernetes while establishing measurable reliability standards.

The role combines software engineering, infrastructure automation, observability, incident management, and resilience engineering.

You will work closely with software, platform, security, QA, and product teams to make systems safer and easier to operate at scale.

A major focus will be reducing operational toil through automation, improving deployment reliability, and turning production incidents into lasting improvements.

You will also contribute to responsible cloud cost management while ensuring performance, availability, and security remain strong.

This role is well suited to a technically strong engineer who enjoys solving complex problems and shaping engineering practices for mission-critical systems.

Accountabilities

    • Define, measure, and continuously improve service reliability through SLIs, SLOs, error budgets, availability targets, and capacity planning.
    • Build and maintain comprehensive observability across infrastructure and applications, including metrics, logs, traces, dashboards, and actionable alerting using tools such as CloudWatch, Prometheus, Grafana, New Relic, ELK, or OpenSearch.
    • Participate in production incident response, troubleshooting, escalation, root-cause analysis, blameless post-incident reviews, and corrective-action tracking.
    • Improve operational excellence through production readiness reviews, runbooks, documentation, change management practices, and engineering standards that reduce risk and improve maintainability.
    • Optimize infrastructure for reliability and performance while managing cloud consumption responsibly and supporting FinOps initiatives.
    • Analyze system performance, resource utilization, latency, throughput, and growth trends to identify bottlenecks and implement scalable solutions proactively.
    • Identify repetitive operational tasks and replace them with reliable automation using Python, Bash, Go, CI/CD tooling, and platform APIs.
    • Design and improve CI/CD pipelines that enable safe, repeatable deployments through automated testing, validation, progressive delivery, rollback strategies, and deployment observability.
    • Apply security and compliance controls across cloud environments, including least-privilege IAM, encryption, network security, secrets management, patching, vulnerability management, and auditability.
    • Design, develop, review, and maintain reusable Terraform modules and infrastructure-as-code patterns for AWS environments.
    • Operate and improve Kubernetes platforms, including EKS clusters, workloads, Helm deployments, autoscaling, upgrades, resource management, networking, storage, and workload resilience.
    • Architect, operate, and optimize AWS services including EC2, S3, RDS, EKS, Lambda, VPC, IAM, Route 53, and related cloud services.
    • Design and validate fault-tolerant architectures, backup strategies, recovery procedures, and disaster recovery capabilities, including reliability testing and failure exercises where appropriate.
    • Partner with development teams to improve application operability, instrumentation, deployment patterns, reliability, and production readiness.
    • Contribute to incident response programs, on-call practices, resilience testing, operational standards, and continuous improvement initiatives across engineering teams.
    • Requirements

      • 5+ years of experience in Site Reliability Engineering, DevOps, platform engineering, cloud infrastructure, or a closely related discipline, with significant production ownership.
      • Strong hands-on experience designing and operating production workloads in AWS, including networking, IAM, compute, storage, databases, DNS, and managed Kubernetes.
      • Advanced Terraform expertise, including reusable modules, remote state, dependency management, environment design, code reviews, and infrastructure lifecycle management.
      • Strong experience with observability and production telemetry, using technologies such as Prometheus, Grafana, CloudWatch, New Relic, ELK, OpenSearch, or comparable platforms.
      • Practical understanding of SRE principles including SLIs, SLOs, error budgets, capacity planning, fault tolerance, graceful degradation, and operational toil reduction.
      • Deep Kubernetes knowledge covering EKS, Helm, workload scheduling, networking, storage, autoscaling, upgrades, troubleshooting, and production operations.
      • Strong Linux systems expertise, with the ability to diagnose issues involving CPU, memory, disk, networking, processes, DNS, and application dependencies.
      • Proficiency in Python, Bash, Go, or another general-purpose programming language used for operational tooling and automation.
      • Experience troubleshooting complex production incidents and contributing to incident response, root-cause analysis, postmortems, and corrective actions.
      • Experience designing or operating CI/CD systems such as GitHub Actions, Jenkins, GitLab CI, Argo CD, or comparable technologies.
      • Working knowledge of cloud security practices, including IAM, encryption, secrets management, network segmentation, vulnerability management, and audit controls.
      • Strong written and verbal communication skills, with the ability to collaborate effectively across technical and business teams.
      • Ability to work independently, take ownership of production systems, and make sound technical decisions in complex environments.
      • AWS certifications such as Solutions Architect Professional or DevOps Engineer Professional are desirable.
      • Experience with GitOps, Argo CD or Flux, formal on-call programs, chaos engineering, resilience testing, service meshes, distributed systems, microservices, RDS, DynamoDB, PostgreSQL, multi-account AWS environments, cloud governance, FinOps, or compliance frameworks such as SOC 2, ISO 27001, or PCI DSS is a plus.
      • Benefits

        • Flexible remote or hybrid work options.
        • Comprehensive health and wellness benefits, including medical, dental, and vision coverage with 100% premium coverage for you.
        • Generous paid time off and paid holidays.
        • MyShare Employee Ownership Program.
        • 401(k) plan with up to a 4% employer match, with full vesting from day one.
        • Opportunities to work alongside industry leaders and experienced engineering professionals.
        • Professional growth opportunities in modern cloud, reliability, and energy technology environments.
        • Opportunity to contribute to advanced energy intelligence and infrastructure supporting a more modern, secure, and resilient grid.
        • Occasional travel of 10% or less may be required.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
Remote • Full-Time • 8+ years exp • PhD • Guadalajara
PowerShell
Python
Databases
Amazon Aurora
DynamoDB
AI/ML
Amazon SageMaker
AWS Bedrock
AWS Bedrock AgentCore
Ray
DevOps
Amazon CloudWatch
Amazon EC2
Amazon ECS
Amazon EKS
Amazon EventBridge
Amazon S3
AWS
AWS Lambda
AWS Step Functions
Azure
CI/CD
Datadog
FinOps
GCP
Git
GitLab
GitLab CI
IAM
Jenkins
JFrog Artifactory
Kubernetes
New Relic
Service Mesh
Splunk
Terraform
Cybersecurity
HIPAA
ISO 27001
PCI DSS
SOC 2
Apply
Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Riga
PowerShell
Python
DevOps
AWS
Azure
Azure AKS
Azure DevOps
Bicep
CI/CD
Configuration Management
Docker
FinOps
GCP
Git
GitLab
Jenkins
Kubernetes
Terraform
Apply
Remote/Hybrid • Full-Time • Buenos Aires
DevOps
AWS
CloudFormation
Terraform
Apply
$87k – $130k per year • In office • Full-Time • 6+ years exp • Murray
Python
TypeScript
JavaScript
Python
Alembic
FastAPI
Pydantic
SQLAlchemy
Databases
PostgreSQL
Redis
AI/ML
Embeddings
LLM
LLM Guardrails
Ollama
RAG
Frontend
React Query
React Router
React.js
Vite
DevOps
AWS
CI/CD
Docker
Docker Compose
IAM
Terraform
Cybersecurity
FedRAMP
NIST 800-53
QA
Pytest
Apply
HLS specialist 24 min ago
In office • Full-Time • Israel
DevOps
AWS
Azure
Apply
$152k – $229k per year • Remote • Full-Time
AI/ML
Human-in-the-Loop
Apply
$162k – $180k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Python
SQL
Databases
Snowflake
AI/ML
dbt
DevOps
AWS
Cybersecurity
HIPAA
Zero Trust
Analytics
ETL/ELT
Apply
$14k – $32k per year (Estimated) • Remote • Full-Time • 2+ years exp • Bachelor's Degree
DevOps
Incident Management
Management
ServiceNow
Apply
$29k – $60k per year (Estimated) • Remote • Full-Time • 8+ years exp
DevOps
Azure
Azure DevOps
Apply
$26k – $69k per year (Estimated) • Remote • Full-Time • 12+ years exp • Bachelor's Degree
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.