368,910open jobs
9,449companies
47,822added this week
Browse all
Salary
$94k – $199k per year (Estimated)
Location
Remote (Canada)
Seniority
Principal · 10+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Herbalife is a global health and wellness community born to support you in living your best life. For over 40 years and in more than 90 countries, we've empowered millions of people to make real changes to their lives with our science-backed products, the support of a coach - what we call an Herbalife Distributor - and the opportunity to build a business. And we're just getting started.

This a Full Remote job, the offer is available from: Canada

Overview:

POSITION SUMMARY:

The SRE Principal Engineer III role is responsible for leading, designing, and implementing robust Site Reliability Engineering (SRE) practices to ensure high availability, scalability, and resilience of critical business systems and applications. The SRE Principal Engineer III will focus on improving system reliability through automation, monitoring, and performance tuning, working closely with development and operations teams to foster a culture of continuous improvement and operational excellence.

The SRE organization spans key disciplines including:

  • SRE Engineering
  • Deployment Automation
  • Incident Response and Postmortem Analysis
  • Observability and Monitoring

Operating in a fully remote capacity, this role will drive the adoption of best practices in multi-cloud and hybrid-cloud platforms, managing services from major cloud providers like Microsoft Azure, Amazon AWS, Oracle OCI, Google GCP, and Alibaba Cloud. The SRE Principal Engineer III will focus on automation, incident management, performance monitoring, and optimizing infrastructure to support scalable, reliable systems. The position will also be responsible for fostering collaboration between development, operations, and security teams to streamline system operations across the organization.

DETAILED RESPONSIBILITIES/DUTIES:

  • Lead the implementation and optimization of SRE practices, ensuring system reliability, performance, and scalability.
  • Architect and maintain automation for infrastructure provisioning, deployment, and incident response.
  • Establish and enforce SLOs (Service Level Objectives) and SLIs (Service Level Indicators) for key services.
  • Collaborate with development teams to design and deliver reliable software systems, ensuring that production environments are optimized for uptime and performance.
  • Create and maintain monitoring, alerting, and observability solutions to provide real-time insights into system health and performance.
  • Respond to production incidents, perform root cause analysis, and implement corrective measures to prevent recurrence.
  • Continuously improve system performance, capacity planning, and reliability through infrastructure tuning and automation.
  • Facilitate post-incident reviews, fostering a blameless culture that focuses on learning from incidents.
  • Collaborate with security teams to ensure infrastructure meets compliance, security standards, and best practices.
  • Foster a collaborative environment across development, operations, and security teams to enhance operational efficiency and knowledge sharing.
  • Drive the adoption of automation tools and frameworks to minimize manual intervention and optimize systems.

Qualifications:

Skills Required:

  • Proven expertise in SRE practices, with a focus on automation, incident management, observability, and infrastructure scalability.
  • Extensive knowledge of cloud platforms (Azure, AWS, GCP, OCI) and hybrid-cloud environments, with a focus on reliability and performance optimization.
  • Experience with automation tools and scripting languages, such as Python, Go, Terraform, or Ansible, for managing infrastructure and incident response.
  • Strong understanding of containerization (Docker, Kubernetes) and orchestration systems.
  • Solid grasp of monitoring and observability tools (Prometheus, Grafana, Dynatrace, Splunk) to ensure real-time system health monitoring.
  • Expertise in capacity planning, performance tuning, and failure management techniques.
  • Strong background in incident management, root cause analysis, and postmortem processes to improve system resilience.
  • Deep understanding of security and compliance requirements, and the ability to ensure production environments meet industry standards.
  • Experience with Agile and DevOps methodologies to ensure fast, reliable delivery of services.

Experience Required:

  • 10+ years of experience in IT, with a focus on SRE, DevOps, or infrastructure engineering roles.
  • Extensive hands-on experience with cloud infrastructure management and automation tools such as Terraform, CloudFormation, or equivalent.
  • Proficiency in scripting and automation languages like Python, Bash, Go, or Ruby for infrastructure automation.
  • Proven experience in managing large-scale systems, ensuring reliability, high availability, and scalability.
  • Expertise in container orchestration technologies, including Kubernetes, OpenShift, and Docker Swarm.
  • Deep knowledge of monitoring and observability platforms (Prometheus, Grafana, ELK, Dynatrace), including experience building and maintaining alerting and dashboard systems.
  • Strong understanding of version control systems and CI/CD practices to optimize code deployment as it relates to infrastructure.
  • Demonstrated ability to optimize performance in multi-cloud and hybrid-cloud environments, ensuring uptime and performance at scale.

Education Required:

  • Bachelor’s degree in Computer Science, Information Technology, or related field, or equivalent experience.

Certificates / Training Preferred:

  • Relevant cloud certifications such as AWS Certified Solutions Architect, Azure Solutions Architect Expert, or Google Cloud Professional Cloud Architect.
  • SRE-related certifications like Certified Kubernetes Administrator (CKA) or Google Professional Cloud DevOps Engineer.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,910 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$40k – $86k per year (Estimated) • Remote/Hybrid • Full-Time • Élancourt
Bash
PowerShell
Python
SQL
Databases
MySQL
PostgreSQL
Redis
DevOps
Amazon CloudWatch
Ansible
AWS
Azure
CentOS Stream
CI/CD
Debian
Docker
GCP
GitLab
GitLab CI
Grafana
Hyper-V
IAM
Jenkins
Kubernetes
Nagios
Prometheus
Proxmox VE
Terraform
Ubuntu
VMWare
Windows Server
Zabbix
Apply
Cloud Scrum Master 11 hours ago
$31k – $64k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • Hyderabad
DevOps
AWS
Azure
GCP
Management
Confluence
Apply
$17k – $43k per year (Estimated) • Remote • Moscow
C++
Java
Python
Java
Maven
DevOps
Ansible
CI/CD
Docker
Git
Graylog
HAProxy
Jenkins
Nginx
Prometheus
Zabbix
Apply
Data Scientist 11 hours ago
$17k – $47k per year (Estimated) • Remote/Hybrid • Full-Time • 1+ year exp • Bachelor's Degree • Pune
Python
SQL
AI/ML
MLFlow
PyTorch
RAG
Scikit-learn
TensorFlow
DevOps
AWS
Azure
GCP
Analytics
Matplotlib
Seaborn
Apply
$18k – $48k per year (Estimated) • In office • Full-Time • 7+ years exp • Mumbai
JavaScript
SQL
TypeScript
Java
Java
Spring Boot
Frontend
Angular
DevOps
AWS
Azure
Kubernetes
Rest API
Apply
See all jobs
This is one of many
368,910 more open roles from verified company boards, updated every day.