663,523open jobs
38,718companies
99,026added this week
Browse all
Salary
$44k – $107k per year (Estimated)
Location
Remote (Brazil)
Seniority
Senior
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is a Belgian recruitment platform built entirely around remote and flexible work, aggregating openings from thousands of employers that allow work from outside an office. Its matching engine ranks roles against a candidate's skills, seniority and stated preferences on location and flexibility, rather than leaving people to filter a keyword search, and it verifies how genuinely remote each posting is. The company also runs an AI screening layer that shortlists applicants for employers, and publishes research and guidance on distributed work practices alongside the job marketplace itself.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a SRE Sênior based in Brazil.

This is a senior Site Reliability Engineering role focused on building and evolving reliable, scalable, secure, and highly automated cloud platforms.

You will work closely with development teams to improve service availability, reduce operational risk, and accelerate software delivery.

The role combines cloud infrastructure, Kubernetes, infrastructure as code, CI/CD, observability, automation, and DevSecOps practices.

You will help establish and measure reliability standards through SLIs, SLOs, SLAs, error budgets, and incident management practices.

You will also play a key role in troubleshooting, root-cause analysis, performance improvement, and preventing recurring incidents.

The environment encourages continuous platform evolution, reusable engineering practices, and the strategic use of AI to increase productivity and operational efficiency.

This opportunity is ideal for an experienced SRE, DevOps, or Platform Engineer who enjoys solving complex infrastructure challenges and enabling engineering teams to deliver with confidence.

Accountabilities:

    • Design, implement, and continuously evolve cloud platforms and environments with a strong focus on reliability, scalability, security, and resilience.

    • Build, maintain, and improve CI/CD pipelines, promoting automation, consistency, and efficient software delivery practices.

    • Provision and manage cloud infrastructure using Terraform and Infrastructure as Code (IaC).

    • Administer, maintain, and evolve Kubernetes and Docker environments.

    • Automate deployment, provisioning, configuration, and operational processes to reduce manual work and operational risk.

    • Implement and continuously improve monitoring, observability, alerting, and reliability capabilities.

    • Define, monitor, and evolve SLIs, SLOs, SLAs, error budgets, and other reliability indicators.

    • Investigate incidents, troubleshoot complex technical problems, and conduct Root Cause Analysis (RCA).

    • Identify preventive measures and improvements to reduce recurring incidents and strengthen platform resilience.

    • Partner with development teams to promote best practices in cloud architecture, observability, performance, security, and reliability.

    • Develop reusable platform standards, components, and engineering patterns.

    • Continuously improve the security, resilience, and operational maturity of cloud environments.

    • Identify opportunities for automation, operational efficiency, and reduction of repetitive manual activities.

    • Contribute to DevSecOps and FinOps practices across the platform.

    • Use AI-powered tools as productivity accelerators for development, automation, troubleshooting, and day-to-day operations.

    • Requirements:

      • Solid professional experience as an SRE, Senior DevOps Engineer, Platform Engineer, or in a closely related role.

      • Hands-on experience managing AWS environments in production.

      • Strong knowledge of Kubernetes and Docker.

      • Proven experience with Terraform and Infrastructure as Code (IaC).

      • Experience building and maintaining CI/CD pipelines using tools such as GitHub Actions, GitLab CI, or Jenkins.

      • Strong Linux administration and automation skills using Bash and/or Python.

      • Practical experience with observability and monitoring platforms such as Prometheus, Grafana, or equivalent solutions.

      • Experience with centralized log management and analysis using ELK, OpenSearch, or similar technologies.

      • Solid understanding of networking, security, cloud architecture, and infrastructure automation.

      • Experience automating infrastructure, deployment, and operational processes.

      • Strong understanding of SRE principles, including availability, reliability, SLIs, SLOs, SLAs, error budgets, and incident management.

      • Proven ability to work closely with development teams to identify and resolve performance, availability, and reliability issues.

      • Mandatory: Practical experience using AI tools for software development and/or operations, such as Kiro, Claude, GitHub Copilot, ChatGPT, Cursor, or equivalent solutions.

      • Strong problem-solving, analytical, communication, and collaboration skills.

      • Ability to work autonomously while partnering effectively with engineering and technical stakeholders.

      • Nice to have: Experience in the financial services industry.

      • Nice to have: Knowledge of Open Finance/Open Banking.

      • Nice to have: Experience with Kafka, Helm, ArgoCD, and service mesh technologies such as Istio.

      • Nice to have: Knowledge of DevSecOps and FinOps practices.

      • Nice to have: Experience designing and developing Internal Developer Platforms (IDPs).

      • Nice to have: Knowledge of GitOps methodologies.

      • Nice to have: Experience with high-availability architectures, disaster recovery, and business/service continuity strategies.

      • Nice to have: Knowledge of capacity planning and performance optimization.

      • Benefits:

        • Remote work model, allowing you to work from home.

        • Opportunity to work on cloud platforms and reliability challenges within the financial services ecosystem.

        • Senior-level technical scope with significant influence over platform architecture, automation, reliability, and engineering practices.

        • Opportunity to work closely with development teams and influence modern DevOps, SRE, DevSecOps, and platform engineering practices.

        • Exposure to modern cloud-native technologies including AWS, Kubernetes, Terraform, CI/CD, observability, and GitOps.

        • Opportunity to leverage cutting-edge AI tools to improve engineering productivity, automation, and troubleshooting.

        • Consulting/contractor engagement model with the opportunity to contribute to strategic technology initiatives.

        • Continuous-learning environment focused on automation, reliability, security, scalability, and operational excellence.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
663,523 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$52k – $130k per year (Estimated) • Remote • Full-Time • 7+ years exp
Python
JavaScript
Java
Kotlin
TypeScript
SQL
Scala
Groovy
Java
Spring Boot
Databases
PostgreSQL
Snowflake
Databricks
Amazon Redshift
AI/ML
AI Agents
LLM
Frontend
React.js
DevOps
AWS
Kubernetes
Incident Management
Apply
$17k – $45k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Python
Java
SQL
C#
Databases
Oracle
QA
Selenium
Robot Framework
Apply
$26k – $47k per year • Remote • Full-Time • 6+ years exp • Bachelor's Degree
Python
SQL
AI/ML
LangGraph
LangChain
LoRA
Fine-tuning
RLHF
Quantization
Scikit-learn
Prompt Engineering
Knowledge Distillation
AI Agents
NLP
PEFT
Llama
Transformers
TensorFlow
PyTorch
LLM
RAG
Semantic Search
OpenAI
LLMOps
TPU
Semantic Search
Multi-Agent Systems
Model Distillation
DevOps
Rest API
GCP
Azure
AWS
Docker
Kubernetes
Vector
Cybersecurity
SOC 2
GDPR
Apply
$17k – $39k per year (Estimated) • Remote/Hybrid • Full-Time • Novosibirsk
Python
SQL
AI/ML
Pandas
NumPy
DevOps
Git
Apply
$178k – $326k per year (Estimated) • In office • 10+ years exp • Bachelor's Degree • Seattle
Python
Java
Scala
Databases
Apache Kafka
AI/ML
Spark
DevOps
AWS
Apply
$52k – $130k per year (Estimated) • Remote • Full-Time • 7+ years exp
Python
JavaScript
Java
Kotlin
TypeScript
SQL
Scala
Groovy
Java
Spring Boot
Databases
PostgreSQL
Snowflake
Databricks
Amazon Redshift
AI/ML
AI Agents
LLM
Frontend
React.js
DevOps
AWS
Kubernetes
Incident Management
Apply
$47k – $96k per year (Estimated) • Remote • Full-Time • 9+ years exp
JavaScript
Java
Java
Spring Boot
Frontend
Vue.js
DevOps
AWS
Apply
Remote • Full-Time • 10+ years exp • Bachelor's Degree
Apply
$30k – $75k per year (Estimated) • Remote • Full-Time • 12+ years exp • Bachelor's Degree
Analytics
ETL/ELT
Apply
Remote • Full-Time • 5+ years exp
SQL
Apply
See all jobs
This is one of many
663,523 more open roles from verified company boards, updated every day.