688,090open jobs
40,138companies
97,702added this week
Browse all
Salary
$101k – $237k per year (Estimated)
Location
Remote/Hybrid (United States)
Seniority
Principal · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is a Belgian recruitment platform built entirely around remote and flexible work, aggregating openings from thousands of employers that allow work from outside an office. Its matching engine ranks roles against a candidate's skills, seniority and stated preferences on location and flexibility, rather than leaving people to filter a keyword search, and it verifies how genuinely remote each posting is. The company also runs an AI screening layer that shortlists applicants for employers, and publishes research and guidance on distributed work practices alongside the job marketplace itself.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineering Team Lead (Principal SRE) based in the United States.

This Principal-level role owns the reliability, availability, and operational health of a global cloud-native AI platform.

You will combine deep hands-on SRE expertise with technical leadership, shaping reliability strategy across critical production services.

The role encompasses SLI/SLO/SLA governance, observability, automation, incident response, production readiness, and high-risk change management.

You will work closely with engineering, DevOps, platform, architecture, and operations teams to embed reliability throughout the software development lifecycle.

As a technical leader without direct reports, you will influence through expertise, mentorship, standards, and informed decision-making.

The environment is distributed and highly technical, with complex systems requiring strong availability, scalability, and operational resilience.

This opportunity is ideal for an experienced SRE professional who enjoys solving challenging infrastructure problems while building sustainable reliability practices.

Accountabilities

    • Provide technical leadership across the Site Reliability Engineering function, helping select, mentor, and develop engineers across multiple locations.

    • Establish technical direction, priorities, engineering standards, and reliability practices while contributing performance and growth feedback to team managers.

    • Own and execute a reliability roadmap covering a 2-3 quarter planning horizon.

    • Define and govern SLI, SLO, and SLA frameworks supporting contracted availability targets of up to 99.95%.

    • Design and maintain a sustainable on-call model while monitoring operational workload, page volume, and team health.

    • Serve as a Tier 2 technical escalation point for major production incidents and collaborate with incident management and operations teams.

    • Promote a blameless postmortem culture and ensure incident reviews result in actionable systemic improvements.

    • Lead Production Readiness and non-functional requirements reviews with development teams.

    • Contribute to root cause analysis and drive reliability improvements resulting from production incidents.

    • Act as an approval authority for high-risk and out-of-window production changes.

    • Define strategic direction for metrics, dashboards, alerting, SLI/SLO monitoring, escalation, and automation.

    • Drive CI/CD automation for service deployments, rollbacks, and operational processes.

    • Partner with DevOps and platform teams to evolve shared infrastructure and reliability capabilities.

    • Work with engineering managers and architects to incorporate reliability principles into the SDLC by default.

    • Participate in architecture reviews and reliability consulting while clearly communicating technical risks and reliability posture to technical and non-technical stakeholders.

    • Requirements

      • 8+ years of hands-on experience in Site Reliability Engineering, DevOps, cloud platforms, or closely related roles, including experience leading a team or owning a technical function.

      • Demonstrated ability to establish technical direction, maintain engineering standards, and influence teams through technical authority, with or without formal management responsibility.

      • Hands-on experience with container orchestration and service technologies such as Kubernetes, Docker, and Istio.

      • Strong experience with public cloud platforms, particularly Azure, with exposure to AWS and Google Cloud.

      • Experience with observability technologies covering metrics, dashboards, and alerting, such as Zabbix, Prometheus, and Grafana.

      • Experience designing and operating CI/CD pipelines and infrastructure-as-code solutions, including technologies such as Terraform and Flux.

      • Proficiency in at least one scripting or programming language, such as Python, Go, or Shell.

      • Strong UNIX/Linux expertise, including system configuration, performance troubleshooting, and networking fundamentals such as Layer 4/5, DNS, HTTP/S, and TLS.

      • Strong understanding of high-availability architecture, including redundancy, failover strategies, and blast-radius management.

      • Excellent written and verbal communication skills in English, with the ability to explain complex technical concepts clearly.

      • Previous SRE leadership experience and experience managing or influencing distributed technical teams are preferred.

      • Experience with log aggregation and analytics platforms such as Loki or Thanos is preferred.

      • Familiarity with ITSM and project management tools such as Jira and Confluence is beneficial.

      • Experience in automotive, embedded systems, or other latency-sensitive production environments is advantageous.

      • Strong collaborative mindset, sound judgment under pressure, and the ability to operate effectively in ambiguous and technically complex environments.

      • Benefits

        • Competitive compensation and benefits package.

        • Annual bonus opportunity.

        • Medical, dental, and vision insurance coverage.

        • Life and disability insurance.

        • Paid time off and paid holidays.

        • Company contribution to an RRSP retirement savings plan.

        • Equity awards for eligible positions and levels.

        • Remote and/or hybrid work options depending on the position and location.

        • Opportunity to work on large-scale cloud-native AI and connected technology platforms.

        • Exposure to distributed engineering teams and complex global production environments.

        • Opportunities to influence technical strategy, reliability standards, and engineering practices at a Principal level.

        • A collaborative environment focused on innovation, technical growth, and continuous improvement.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
688,090 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
DevOps Engineer 2 hours ago
$28k – $60k per year (Estimated) • In office • 5+ years exp • Bengaluru
Python
PowerShell
Bash
Databases
Databricks
Delta Lake
Apache Kafka
AI/ML
Spark
DevOps
Terraform
Ansible
GCP
Azure DevOps
GitHub Actions
CloudFormation
Prometheus
Azure
CI/CD
Jenkins
AWS
Docker
Kubernetes
Grafana
Configuration Management
Amazon EKS
Google GKE
Azure AKS
Amazon CloudWatch
Linux
Unix
Apply
$126k – $189k per year • Remote/Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • Irving
Python
Java
Java
Spring Boot
AI/ML
AI Agents
DevOps
GCP
OpenTelemetry
CI/CD
Git
Docker
Kubernetes
Self-Healing
AIOps
Incident Management
Apply
Platform Architect 2 hours ago
$46k – $96k per year (Estimated) • In office • 10+ years exp • Delhi
Python
Java
Databases
PostgreSQL
Apache Kafka
DevOps
gRPC
GCP
Azure
AWS
Kubernetes
Chaos Engineering
Progressive Delivery
Apply
AWS Devops 2 hours ago
In office • Full-Time • Hyderabad
Python
TypeScript
Databases
PostgreSQL
Redis
RabbitMQ
OpenSearch
DevOps
Terraform
Ansible
GitHub Actions
Istio
Fluent Bit
Prometheus
HAProxy
CI/CD
ArgoCD
Jenkins
Git
AWS
Kubernetes
Nginx
Grafana
Blue-Green Deployment
Platform Engineering
Service Mesh
GitHub
Cybersecurity
SonarQube
Trivy
Apply
$36k – $73k per year (Estimated) • Remote/Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • Gurgaon
Python
SAS
DevOps
Unix
Analytics
Tableau
Apply
$115k – $216k per year (Estimated) • Remote • Full-Time • 6+ years exp
AI/ML
Claude
ChatGPT
Stable Diffusion
AI Agents
Midjourney
ComfyUI
LLM
Runway
Multi-Agent Systems
Apply
$400k – $680k per year • Equity • Remote • Full-Time
Python
Java
DevOps
CI/CD
AWS
IAM
Cybersecurity
Okta
Auth0
LDAP
Apply
$87k – $102k per year • Remote • Full-Time • 4+ years exp • Bachelor's Degree
Python
SQL
AI/ML
Claude
ChatGPT
Analytics
Tableau
Microsoft Excel
Apply
Remote • Full-Time • Bachelor's Degree
Marketing
Salesforce
Apply
$69k – $98k per year • Remote • Full-Time • 4+ years exp • Bachelor's Degree
Analytics
Tableau
Power BI
Alteryx
Apply
See all jobs
This is one of many
688,090 more open roles from verified company boards, updated every day.