664,245open jobs
38,805companies
99,560added this week
Browse all
Salary
$177k – $240k per year
Location
Remote (Brazil)
Seniority
Staff · 10+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is a Belgian recruitment platform built entirely around remote and flexible work, aggregating openings from thousands of employers that allow work from outside an office. Its matching engine ranks roles against a candidate's skills, seniority and stated preferences on location and flexibility, rather than leaving people to filter a keyword search, and it verifies how genuinely remote each posting is. The company also runs an AI screening layer that shortlists applicants for employers, and publishes research and guidance on distributed work practices alongside the job marketplace itself.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Site Reliability Engineer based in Brazil.

This is a high-impact reliability leadership role within a fully remote engineering organization operating globally.

You will be the first dedicated SRE, helping establish reliability practices across multiple engineering teams and critical production systems.

The role combines hands-on engineering with organization-wide influence, covering observability, incident response, operational readiness, and resilience.

You will work closely with engineering leadership, infrastructure specialists, architects, and product teams to make reliability measurable and actionable.

A major focus will be embedding SRE principles into engineering culture rather than simply owning individual services.

You will also help shape how AI is used for incident investigation, operational tooling, observability, and safe system operations.

The position offers substantial autonomy to define standards, coach engineers, and build practices that scale with the organization.

Accountabilities

    • Define and implement SLIs and SLOs for critical production request paths, ensuring reliability objectives are visible, measurable, reviewed, and connected to engineering decisions.

    • Introduce and champion error budgets as a practical framework for balancing reliability investments with product and feature delivery.

    • Establish and maintain the reliability metrics used by engineering leadership to evaluate progress and identify areas requiring investment.

    • Strengthen the complete incident management lifecycle, including detection, response, communication, escalation, postmortems, and follow-up actions.

    • Improve alert quality, anomaly detection, escalation processes, and shared operational tooling in collaboration with infrastructure teams.

    • Lead reliability assessments for high-risk changes and new services, covering production readiness, capacity, failure modes, rollback strategies, and operational risks.

    • Introduce deliberate failure testing, game days, and chaos exercises to identify weaknesses and validate safe operational limits before incidents occur.

    • Work directly with engineering teams on complex reliability challenges through focused engagements, leaving behind stronger practices and clear ownership.

    • Coach Staff and Lead engineers to become reliability advocates within their respective teams and help establish distributed SRE ownership.

    • Develop lightweight, repeatable operational standards covering production readiness, on-call practices, runbooks, change safety, and service operability.

    • Partner with architects and technical leads to ensure reliability and failure tolerance are incorporated into system design rather than addressed after deployment.

    • Remain hands-on during production incidents and investigations, building tooling, dashboards, automation, and reference implementations where appropriate.

    • Promote effective use of AI for incident investigation, telemetry analysis, postmortem development, runbook creation, observability, and reliability tooling.

    • Help structure operational data, alerts, dashboards, and runbooks so that both engineers and AI agents can safely interpret and act on production signals.

    • Contribute production fixes and improvements directly through code and infrastructure changes rather than limiting the role to recommendations and reviews.

    • Requirements

      • 10+ years of engineering experience, including at least 3 years in SRE, production engineering, or a reliability-focused Staff Engineer role operating across multiple teams.

      • Demonstrated experience owning reliability at a platform or organizational level rather than only for an individual service.

      • Deep practical experience designing and implementing SLIs, SLOs, and error budgets, including successfully driving adoption across product and engineering teams.

      • Strong incident leadership experience, including managing high-severity, customer-facing incidents and leading effective postmortems that result in measurable improvements.

      • Advanced understanding of distributed-system failure modes, including database and cache saturation, cascading failures, retry storms, capacity constraints, graceful degradation, and load shedding.

      • Strong hands-on experience with Kubernetes, AWS, and modern observability platforms such as Datadog or comparable technologies.

      • Ability to read and write production code in Go, TypeScript, or a similar language, as well as work with infrastructure as code.

      • Demonstrated ability to influence teams without direct authority and successfully change engineering practices across an organization.

      • Strong coaching and mentoring skills, with evidence of developing engineers into effective reliability owners.

      • Exceptional written and verbal communication skills, with the ability to clearly communicate incidents, risks, technical trade-offs, and reliability priorities to both engineers and executives.

      • Strong preference for asynchronous, documented decision-making and clear technical communication.

      • Practical experience using AI tools for incident investigation, telemetry analysis, runbook and postmortem development, and engineering tooling.

      • Understanding of how operational data, alerts, dashboards, and runbooks should be structured to support safe AI-assisted diagnosis and operations.

      • Pragmatic approach to reliability, with the ability to balance operational risk, engineering investment, delivery speed, and business priorities.

      • Experience in fraud detection, identity, payments, or other real-time and adversarial environments is an asset.

      • Experience with multi-region architectures, cell-based architectures, or failure-isolation strategies is a plus.

      • Experience operating Elasticsearch, Redis, DynamoDB, or Kafka at scale and understanding their failure modes is beneficial.

      • Familiarity with FinOps and cloud infrastructure cost-versus-reliability trade-offs is an advantage.

      • Must be authorized to work from the hiring location; visa sponsorship is not provided.

      • Benefits

        • Fully remote working environment.

        • Opportunity to become the first dedicated Site Reliability Engineer and establish organization-wide reliability practices.

        • High level of autonomy and direct influence over engineering standards, operational practices, and platform reliability.

        • Opportunity to work across multiple engineering teams and critical production systems.

        • Close collaboration with engineering leadership, architects, infrastructure teams, and technical leads.

        • Opportunity to shape AI-assisted reliability practices and the future of production operations.

        • Strong focus on professional growth, technical leadership, coaching, and knowledge sharing.

        • Inclusive, globally distributed engineering environment that values diverse perspectives and backgrounds.

        • For US-based employees, the stated cash compensation range is $177,000-$240,000 USD, with actual offers varying according to factors such as experience, skills, education, certifications, and market conditions. Compensation may differ for other hiring locations.

        • Remote work eligibility is subject to applicable regulatory and security requirements in the candidate's location.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
664,245 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$177k – $240k per year • Remote • Full-Time • 10+ years exp
TypeScript
Databases
Redis
DynamoDB
ElasticSearch
Apache Kafka
AI/ML
AI Agents
DevOps
Datadog
AWS
Kubernetes
FinOps
Incident Management
Apply
$24k – $52k per year (Estimated) • Remote • Full-Time • 4+ years exp • Bachelor's Degree • Bengaluru
Python
DevOps
Splunk
Terraform
Helm
Istio
Consul
Datadog
Linkerd
Prometheus
GitLab CI
CI/CD
GitOps
ArgoCD
Jenkins
AWS
Kubernetes
Grafana
Blue-Green Deployment
Chaos Engineering
Service Mesh
Amazon EKS
Progressive Delivery
FinOps
Cybersecurity
PCI DSS
SOC 2
Zero Trust
Apply
$33k – $60k per year (Estimated) • Remote • 5+ years exp • Tyumen
Python
Go
SQL
Python
Django
Databases
PostgreSQL
RabbitMQ
ElasticSearch
Apache Kafka
AI/ML
Copilot
Cursor
Airflow
Claude Code
AI Agents
DevOps
Rest API
Prometheus
WebSockets
Yandex Cloud
GitLab CI
CI/CD
Docker
Kubernetes
Grafana
Cybersecurity
Keycloak
Analytics
ETL/ELT
Management
Scrum
Kanban
Apply
Team lead (C#/.net) 3 hours ago
up to $94k per year (net) • In office • 3+ years exp • Moscow
C#
C#
.NET
Databases
PostgreSQL
Apache Kafka
AI/ML
AI Agents
DevOps
OpenShift
Istio
Kubernetes
Apply
Remote/Hybrid • Full-Time • Bachelor's Degree • Bengaluru • Pune
Python
Go
Databases
MySQL
PostgreSQL
Apache Kafka
AI/ML
Fine-tuning
AI Agents
LLM
DevOps
Kubernetes
Management
Agile
Apply
$111k – $211k per year (Estimated) • Remote • Full-Time
AI/ML
AI Agents
Apply
$67k – $140k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Python
Ruby
PowerShell
Bash
DevOps
Windows Server
AWS
Nginx
Amazon CloudWatch
Management
Smartsheet
Apply
Head of Product 3 hours ago
$43k – $90k per year (Estimated) • Remote • Full-Time • 10+ years exp
Management
ClickUp
Jira
Apply
$45k – $102k per year (Estimated) • Remote • Full-Time
AI/ML
Physical AI
Apply
$177k – $240k per year • Remote • Full-Time • 10+ years exp
TypeScript
Databases
Redis
DynamoDB
ElasticSearch
Apache Kafka
AI/ML
AI Agents
DevOps
Datadog
AWS
Kubernetes
FinOps
Incident Management
Apply
See all jobs
This is one of many
664,245 more open roles from verified company boards, updated every day.