368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$42k – $105k per year (Estimated)
Location
Remote (Brazil)
Seniority
Senior · 5+ years exp
Overview
Company
Impact
Profile match
Oowlish is a nearshore technology services provider headquartered in Recife, Brazil, and founded in 2017. The company offers a range of services including staff augmentation, AI transformation, digital product development, and specialized e-commerce solutions for platforms like Shopify. It operates primarily by connecting North American businesses with technical talent from Latin America, maintaining a focus on industries such as pet tech, fintech, and edtech.

About the Role:

We are looking for an experienced Senior Site Reliability Engineer (SRE) to own the reliability, availability, and operational excellence of business-critical production systems.

This is a dedicated Site Reliability Engineering role-not a general DevOps or Infrastructure position. You will define how reliability is measured, lead incident response during production outages, drive observability strategy, and continuously improve operational practices across high-availability environments.

The ideal candidate has hands-on experience managing SLOs, leading major incidents, improving on-call operations, and building a strong reliability culture through automation, observability, and continuous improvement.

Responsibilities:

  • Define, implement, and continuously improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
  • Develop and maintain observability strategies, including monitoring, logging, tracing, and alerting.
  • Own observability configuration, instrumentation, and alert optimization.
  • Lead Incident Command during production incidents and coordinate cross-functional response efforts.
  • Drive blameless postmortems and ensure corrective actions are completed.
  • Own and continuously improve the on-call program, including rotations, escalation policies, runbooks, and alert tuning.
  • Establish production readiness standards for new services.
  • Partner with engineering teams on capacity planning, scalability, and disaster recovery initiatives.
  • Automate operational processes and reliability improvements using software engineering best practices.
  • Continuously improve system reliability, availability, and operational efficiency.

Requirements:

  • 5+ years of experience in Site Reliability Engineering, Production Engineering, Reliability Engineering, or similar roles.
  • Proven experience operating production systems in high-availability environments.
  • Hands-on experience defining and managing SLOs, SLIs, and Error Budgets.
  • Experience leading production incident response and Incident Command.
  • Strong observability and monitoring experience.
  • Strong software engineering skills using Python, Go, or TypeScript.
  • Experience working with cloud platforms.
  • Strong written and verbal English communication skills.

Must have:

  • Proven Site Reliability Engineering experience.
  • Experience defining and managing:
    • Service Level Indicators (SLIs)
    • Service Level Objectives (SLOs)
    • Error Budgets
    • Experience leading Incident Command during major production incidents.
    • Experience conducting blameless postmortems and driving follow-up actions.
    • Experience designing, maintaining, and improving on-call programs.
    • Experience developing runbooks and escalation policies.
    • Strong observability experience, including:
      • Monitoring
      • Logging
      • Alerting
      • Distributed Tracing
      • Experience tuning alerts to reduce operational noise.
      • Strong automation skills using Python, Go, or TypeScript.
      • Experience supporting mission-critical production systems.
      • Experience working in high-availability production environments.

Nice to have:

  • Experience with Datadog.
  • Experience with AWS.
  • Experience with Heroku.
  • Experience working in regulated industries (Healthcare, HIPAA, Financial Services, etc.).
  • Experience establishing or maturing an SRE practice.
  • Capacity planning experience.
  • Disaster recovery planning and execution.
  • Experience with Kubernetes.
  • Experience with PostgreSQL or SQL Server.
  • Experience supporting modern TypeScript-based applications.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Brasília
IT Administrator 3 hours ago
$83k – $113k per year • In office • Full-Time
Node JS
Python
JavaScript
DevOps
AWS
GCP
GitHub
Cybersecurity
Okta
SOC 2
Management
Confluence
Jira
Slack
Apply
$63k – $86k per year • Equity • In office • Full-Time • Bachelor's Degree • Austin
Python
SQL
Databases
Databricks
Analytics
Tableau
Apply
Data Scientist 3 hours ago
$113k – $188k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Arlington • Washington
Python
Databases
Databricks
DevOps
AWS
Azure
Analytics
ETL/ELT
Power BI
Apply
$61k – $148k per year (Estimated) • In office • Full-Time • 6+ years exp • Bachelor's Degree • Erlangen
C++
Python
SystemVerilog
VHDL
DevOps
CI/CD
Apply
$173k – $260k per year • In office • Full-Time • PhD • San Francisco
JavaScript
Node JS
Python
Python
Celery
Django
Flask
Databases
RabbitMQ
Redis
AI/ML
Agentforce
AI Agents
DevOps
Akamai
AWS
CI/CD
Cloudflare
CloudFormation
Helm
Jenkins
Kubernetes
Spinnaker
Terraform
Marketing
Salesforce
Apply
$55k – $119k per year (Estimated) • Remote • 8+ years exp • Rio de Janeiro
DevOps
AWS
OpenTelemetry
Terraform
Apply
$29k – $86k per year (Estimated) • Remote • 4+ years exp • Rio de Janeiro
DevOps
AWS
CI/CD
OpenTelemetry
Terraform
Management
Jira
Apply
$29k – $86k per year (Estimated) • Remote • 4+ years exp • Brasília
DevOps
AWS
CI/CD
OpenTelemetry
Terraform
Management
Jira
Apply
$55k – $119k per year (Estimated) • Remote • 8+ years exp • Brasília
DevOps
AWS
OpenTelemetry
Terraform
Apply
$50k – $132k per year (Estimated) • Remote • Brasília
Node JS
JavaScript
Node JS
Mongoose
Nest.JS
Databases
Redis
Supabase
Frontend
Vue.js
DevOps
AWS
CI/CD
GitHub Actions
Terraform
Vercel
GitHub
Apply
$27k – $92k per year (Estimated) • Remote • Full-Time • 4+ years exp • Brasília
PHP
JavaScript
PHP
Eloquent
Laravel
Livewire
Pest
PHPUnit
Frontend
Inertia.js
React.js
Tailwind CSS
Vue.js
QA
Playwright
Apply
$40k – $115k per year (Estimated) • Remote • Full-Time • 5+ years exp • Brasília
JavaScript
Python
TypeScript
Python
Django
FastAPI
Flask
Databases
PostgreSQL
Frontend
React.js
Apply
$40k – $115k per year (Estimated) • Remote • Full-Time • 5+ years exp • Brasília
JavaScript
Python
TypeScript
Python
Django
FastAPI
Flask
Databases
PostgreSQL
Frontend
React.js
Apply
$63k – $130k per year (Estimated) • Remote • Full-Time • Brasília
DevOps
Rest API
Apply
$29k – $86k per year (Estimated) • Remote • 4+ years exp • Brasília
DevOps
AWS
CI/CD
OpenTelemetry
Terraform
Management
Jira
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.