368,530open jobs
9,432companies
50,439added this week
Browse all
Location
Remote (Brazil)
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is an AI-powered job platform focused on remote and flexible work. It matches candidates with relevant roles using skills and preference-based algorithms, and also offers career coaching and job-search guidance.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a SRE Engineer based in Brazil.

This is an opportunity for an SRE professional to strengthen the reliability, resilience, and performance of critical digital environments.

You will work across cloud infrastructure, Kubernetes, observability, automation, and incident management to keep systems stable and highly available.

The role combines proactive engineering with hands-on troubleshooting, helping identify risks and eliminate recurring operational issues.

You will define and monitor reliability metrics such as SLI, SLO, SLA, MTTR, and MTTD to drive measurable improvements.

You will collaborate with multidisciplinary teams to embed reliability and observability into solutions from the design stage onward.

Automation, Infrastructure as Code, capacity planning, and continuous improvement will be central to reducing operational toil and improving scalability.

The environment values technical ownership, collaboration, data-driven decisions, and a strong culture of engineering excellence.

Accountabilities:

    • Define, implement, and monitor SLIs, SLOs, SLAs, MTTR, and MTTD, establishing measurable reliability objectives.
    • Implement and evolve observability solutions, including monitoring, alerting, dashboards, and APM.
    • Monitor system latency, traffic, errors, saturation, availability, and overall application and infrastructure performance.
    • Prevent, investigate, and resolve incidents, contributing to effective incident response and service restoration.
    • Conduct root-cause analyses and establish corrective and preventive actions to avoid recurring incidents.
    • Identify infrastructure risks, bottlenecks, single points of failure, and opportunities to strengthen system resilience.
    • Support the design and evolution of resilient, scalable, and highly available solutions.
    • Automate operational activities and reduce repetitive manual work and operational toil.
    • Operate and continuously improve Kubernetes and Docker environments.
    • Support capacity planning, business continuity, and disaster-recovery strategies.
    • Participate in deployments and help stabilize applications and environments after releases.
    • Partner with development, infrastructure, security, and other teams to incorporate reliability practices from the earliest stages of solution design.
    • Create and maintain operational dashboards, alerts, procedures, runbooks, and technical documentation.
    • Promote a culture centered on reliability, observability, automation, and continuous improvement.
    • Requirements:

      • Professional experience working as a Site Reliability Engineer (SRE) or in an equivalent reliability, DevOps, or infrastructure engineering role.
      • Practical experience with cloud environments, particularly GCP, AWS, and/or Azure.
      • Solid knowledge of Kubernetes and Docker.
      • Experience with observability, monitoring, alerting, and APM solutions.
      • Strong understanding of SRE concepts and metrics, including SLI, SLO, SLA, MTTR, MTTD, and error budgets.
      • Experience managing, investigating, and resolving production incidents.
      • Strong troubleshooting skills across applications and infrastructure.
      • Experience administering Linux environments.
      • Knowledge of networking, security, performance optimization, scalability, and high availability.
      • Experience with automation and Infrastructure as Code (IaC).
      • Experience working with CI/CD pipelines and modern software delivery practices.
      • Strong analytical and problem-solving capabilities, with a proactive approach to preventing issues before they affect production.
      • Excellent communication skills and the ability to collaborate effectively with multidisciplinary engineering teams.
      • Differentials: experience with GKE, EKS, or AKS; Dynatrace, Datadog, Grafana, Prometheus, ELK, Elasticsearch, or Kibana; Terraform and Ansible; distributed and mission-critical systems; regulated or financial environments; cloud capacity and cost optimization; disaster recovery and business continuity; and cloud, Kubernetes, or SRE certifications.
      • Benefits:

        • Meal allowance (Vale Refeição).
        • Food allowance (Vale Alimentação).
        • Home office allowance.
        • Medical insurance.
        • Dental insurance.
        • Life insurance.
        • Birthday day off.
        • TotalPass / Wellhub wellness benefit.
        • Access to the Boon Saúde health platform.
        • Discounts and partnerships with businesses and educational institutions.
        • Welcome kit.
        • Structured onboarding program.
        • Access to continuous learning through Verity Learning.
        • Internal initiatives focused on knowledge sharing and professional development.
        • Programs and initiatives supporting employee well-being and connection.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$11k – $14k per year • In office • 2+ years exp • Almaty
Bash
Python
Databases
Apache Kafka
PostgreSQL
RabbitMQ
Redis
DevOps
Ansible
ArgoCD
AWS
Azure
CentOS Stream
CI/CD
Docker
FluxCD
Git
GitHub Actions
GitLab CI
GitOps
Grafana
HAProxy
Hyper-V
Jenkins
Kubernetes
Nginx
OpenStack
Prometheus
Proxmox VE
Terraform
Ubuntu
VMWare
Yandex Cloud
Zabbix
GitHub
GitLab
Apply
$72k – $180k per year (Estimated) • In office • 5+ years exp • São Paulo
Go
Python
AI/ML
Reinforcement Learning
AI Agents
Edge AI
DevOps
Ansible
AWS
Azure
Chef
CI/CD
Docker
GCP
GitHub Actions
Google GKE
Jenkins
Kubernetes
Platform Engineering
Puppet
Terraform
GitHub
Cybersecurity
GDPR
Apply
$133k – $181k per year • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Cooper
Bash
PowerShell
Python
DevOps
Ansible
Azure
Configuration Management
Docker
Kubernetes
OpenShift
Ubuntu
Windows Server
Cybersecurity
CIS Benchmarks
Defense in Depth
Nessus
Qualys Cloud Platform
Zero Trust
Apply
$143k – $184k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • Cooper
DevOps
Azure
IAM
Cybersecurity
FedRAMP
Microsoft Defender
Nessus
NIST 800-53
Apply
$153k – $207k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • Cooper
PowerShell
DevOps
Azure
Docker
Kubernetes
Windows Server
IAM
Cybersecurity
Keycloak
Microsoft Entra ID
Okta
Zero Trust
Apply
$126k – $201k per year • Equity • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Analytics
A/B Testing
Apply
$84k – $166k per year (Estimated) • Remote • Full-Time • 7+ years exp • Bachelor's Degree
SQL
Apply
$80k – $190k per year • Remote • Full-Time • 2+ years exp
Apply
$134k – $223k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Bash
Python
AI/ML
Claude
Claude Code
Copilot
OpenAI Codex
DevOps
Azure
Azure DevOps
CI/CD
Gerrit
Git
Jenkins
KVM
QEMU
RTOS
VMWare
Xen
Cybersecurity
Tcpdump
Wireshark
IoT
FreeRTOS
Management
Confluence
Jira
Apply
$165k – $301k per year (Estimated) • Equity • Remote • Full-Time • 12+ years exp
AI/ML
AI Agents
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.