368,530open jobs
9,432companies
50,439added this week
Browse all
Location
Remote (Brazil)
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is an AI-powered job platform focused on remote and flexible work. It matches candidates with relevant roles using skills and preference-based algorithms, and also offers career coaching and job-search guidance.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a SRE Engineer based in Brazil.

As an SRE Engineer, you will play a key role in ensuring the reliability, scalability, performance, and resilience of critical digital environments.

You will work across cloud infrastructure, Kubernetes, observability, automation, and incident management to keep systems stable and highly available.

The role combines proactive engineering with hands-on troubleshooting, helping teams identify risks before they become production issues.

You will establish and monitor reliability metrics such as SLIs, SLOs, SLAs, MTTR, and MTTD to continuously improve operational performance.

You’ll collaborate closely with multidisciplinary teams to embed reliability practices throughout the software development lifecycle.

Automation and Infrastructure as Code will be central to reducing manual work and creating more efficient, consistent operations.

This is an opportunity to contribute to a culture of continuous improvement while working on modern, cloud-based and distributed technology environments.

Accountabilities:

    • Define, monitor, and continuously improve reliability indicators, including SLIs, SLOs, SLAs, MTTR, MTTD, and error budgets.
    • Implement and evolve observability, monitoring, alerting, and APM solutions across applications and infrastructure.
    • Monitor latency, traffic, errors, saturation, availability, and overall system performance.
    • Prevent, investigate, and resolve incidents, minimizing their impact on users and business operations.
    • Conduct root cause analyses and define corrective and preventive actions to avoid recurring incidents.
    • Identify operational risks, bottlenecks, single points of failure, and opportunities to strengthen system resilience.
    • Support the design and evolution of highly available, scalable, resilient, and fault-tolerant solutions.
    • Automate operational activities and reduce repetitive manual tasks through scripting, automation, and Infrastructure as Code.
    • Operate and continuously improve Kubernetes and Docker environments.
    • Support capacity planning, business continuity, disaster recovery, and cloud cost optimization initiatives.
    • Participate in deployments and contribute to application stabilization following releases.
    • Collaborate with engineering, development, infrastructure, and other technical teams to incorporate reliability from the earliest stages of solution design.
    • Create and maintain operational dashboards, alerts, procedures, runbooks, and technical documentation.
    • Promote a culture centered on reliability, observability, automation, prevention, and continuous improvement.
    • Requirements:

      • Proven professional experience as a Site Reliability Engineer, SRE, or in an equivalent reliability/platform engineering role.
      • Practical experience with cloud environments, using one or more of GCP, AWS, or Azure.
      • Hands-on knowledge of Kubernetes and Docker.
      • Experience implementing and managing observability, monitoring, alerting, and APM solutions.
      • Strong understanding of SRE concepts and metrics, including SLI, SLO, SLA, MTTR, MTTD, and error budgets.
      • Experience managing, investigating, troubleshooting, and resolving production incidents.
      • Knowledge of application and infrastructure troubleshooting in complex environments.
      • Experience administering Linux environments.
      • Understanding of networking, security, performance, scalability, and high availability.
      • Experience with automation and Infrastructure as Code practices.
      • Hands-on experience with CI/CD pipelines and modern software delivery practices.
      • Strong communication and collaboration skills, with the ability to work effectively across multidisciplinary teams.
      • Analytical, proactive, collaborative, and prevention-oriented mindset.
      • Experience with GKE, EKS, or AKS is a plus.
      • Knowledge of Dynatrace, Datadog, Grafana, Prometheus, or comparable observability platforms is a plus.
      • Experience with ELK Stack, Elasticsearch, and Kibana is desirable.
      • Knowledge of Terraform and Ansible is desirable.
      • Experience supporting critical systems and distributed architectures is an advantage.
      • Experience in financial institutions or other regulated environments is a plus.
      • Experience with cloud capacity management and cost optimization is desirable.
      • Knowledge of disaster recovery and business continuity practices is beneficial.
      • Cloud, Kubernetes, or SRE certifications are considered a plus.
      • Benefits:

        • Meal and food allowance.
        • Home office allowance.
        • Medical insurance.
        • Dental insurance.
        • Life insurance.
        • Birthday Day Off.
        • TotalPass / Wellhub access.
        • Health and wellness support through the Boon Saúde app.
        • Discounts and partnerships with a variety of establishments.
        • Partnerships with educational institutions and other services.
        • Welcome kit.
        • Structured onboarding program.
        • Access to continuous learning and professional development initiatives.
        • Dedicated learning and knowledge-sharing programs.
        • Employee support and engagement initiatives.
        • Fully remote work opportunity.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$105k – $252k per year • Remote • Full-Time • 18+ years exp • Bachelor's Degree
Python
Java
Java
Gradle
DevOps
Ansible
AWS
CI/CD
CloudFormation
Configuration Management
Docker
GitHub Actions
GitLab CI
Helm
Jenkins
Kubernetes
Platform Engineering
Terraform
GitHub
GitLab
Cybersecurity
Sonatype Nexus IQ
Management
Confluence
Jira
Apply
$54k – $175k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$35k – $113k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$35k – $116k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$133k – $161k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Westminster
C++
Python
DevOps
CI/CD
SpaceTech
NASA cFS
Apply
$73k – $154k per year (Estimated) • Remote • Full-Time • 4+ years exp • Bachelor's Degree
DevOps
VMWare
Apply
$83k – $139k per year • Remote • Full-Time • 3+ years exp • Bachelor's Degree
Analytics
Power BI
Tableau
Apply
$191k – $267k per year • Equity • Remote • Full-Time • 4+ years exp • Master's Degree
Python
SQL
AI/ML
Anomaly Detection
Apply
$217k – $304k per year • Equity • Remote • Full-Time • 8+ years exp
Go
Databases
Apache Kafka
ClickHouse
Google BigQuery
AI/ML
Flink
Recommender Systems
DevOps
Incident Management
Kubernetes
Apply
$152k – $239k per year • Remote • Full-Time • 8+ years exp
SQL
Databases
Snowflake
AI/ML
LLM
Model Context Protocol
Analytics
A/B Testing
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.