685,335open jobs
39,864companies
96,579added this week
Browse all
Salary
$43k – $106k per year (Estimated)
Location
Remote (Brazil)
Seniority
Senior · 12+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is a Belgian recruitment platform built entirely around remote and flexible work, aggregating openings from thousands of employers that allow work from outside an office. Its matching engine ranks roles against a candidate's skills, seniority and stated preferences on location and flexibility, rather than leaving people to filter a keyword search, and it verifies how genuinely remote each posting is. The company also runs an AI screening layer that shortlists applicants for employers, and publishes research and guidance on distributed work practices alongside the job marketplace itself.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an Expert Observability Engineer based in Brazil.

This role provides senior technical leadership across enterprise observability, reliability engineering, and cloud-native monitoring environments.

You will design and govern unified observability strategies spanning metrics, logs, traces, and events across complex technology landscapes.

The position combines architecture, automation, SRE practices, incident management, and platform engineering to improve service reliability.

You will lead innovative implementations, modernize legacy monitoring capabilities, and establish secure, scalable, production-ready observability patterns.

The role also involves close collaboration with global engineering and operations teams during major incidents, transitions, and knowledge-transfer programs.

Automation and infrastructure-as-code will be central to reducing operational effort, alert noise, and mean time to detect and resolve issues.

This is an opportunity to influence enterprise technology roadmaps while mentoring technical teams in a remote-first environment.

Accountabilities:

    • Architect and govern a unified enterprise observability framework covering metrics, logs, traces, and events.
    • Design observability solutions using platforms and technologies such as IBM Instana, Grafana, OpenTelemetry, Telegraf, InfluxDB, and Prometheus.
    • Lead First-of-a-Kind (FOAK) implementations, evaluating emerging observability technologies and converting them into secure, repeatable production patterns.
    • Define enterprise standards for telemetry pipelines, data retention, high-cardinality management, and observability cost optimization.
    • Establish and govern Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
    • Act as a senior technical escalation point for critical operational issues and major incidents.
    • Lead P1/P2 incident war rooms and drive evidence-based Root Cause Analysis (RCA).
    • Reduce Mean Time to Detect (MTTD), Mean Time to Resolve (MTTR), and alert noise through event correlation, dynamic thresholds, and dependency mapping.
    • Drive Observability-as-Code and infrastructure automation using Ansible, Terraform, Python, and GitOps practices.
    • Automate deployment, configuration, monitoring, and self-healing workflows for agents and telemetry collectors.
    • Integrate observability platforms with ITSM tools, ServiceNow, Netcool, and CI/CD pipelines.
    • Design deep observability capabilities across Docker, Kubernetes, OpenShift, microservices, and multi-cloud environments.
    • Correlate application performance telemetry with Kubernetes control planes, pods, nodes, and infrastructure dependencies.
    • Establish secure-by-design telemetry pipelines incorporating RBAC, TLS, secrets management, and image scanning.
    • Lead complex knowledge-transfer programs, vendor transitions, and operational-readiness handovers for global 24x7 teams.
    • Mentor cross-functional engineering and operations teams on observability, SRE, automation, and reliability practices.
    • Contribute to enterprise technology roadmaps and influence architectural decisions across observability and platform engineering.
    • Modernize legacy monitoring environments by introducing proactive, automated, and scalable observability capabilities.
    • Requirements:

      • 12+ years of total IT experience, with at least 5-7 years operating as a Lead Architect, SRE, Principal Observability Engineer, or equivalent senior technical role in a large enterprise environment.
      • Strong hands-on experience with observability and APM technologies, including IBM Instana, Grafana Enterprise and Alloy, Prometheus, OpenTelemetry, Telegraf, and InfluxDB.
      • Experience with traditional monitoring platforms such as SolarWinds, Netcool, Elastic, and Splunk.
      • Proven expertise with Kubernetes, Docker, OpenShift, and cloud environments across AWS, Azure, and/or GCP.
      • Strong infrastructure knowledge across Linux/RHEL, Windows Server, VMware, Citrix VDI, load balancers, and edge proxies.
      • Hands-on experience with Ansible, Terraform, Python, Bash, webhooks, and CI/CD platforms such as GitHub Actions, GitLab, or Jenkins.
      • Experience integrating monitoring and observability platforms with ServiceNow and other ITSM and operational tooling.
      • Strong understanding of ITIL 4 principles and advanced Major Incident Management.
      • Proven track record of migrating organizations from legacy monitoring solutions to proactive and automated observability platforms.
      • Experience designing scalable and secure telemetry pipelines and time-series database environments.
      • Strong understanding of SRE concepts, including SLIs, SLOs, error budgets, reliability engineering, and incident response.
      • Experience implementing observability across distributed systems, microservices, Kubernetes, and multi-cloud architectures.
      • Demonstrated experience leading FOAK technology implementations and complex vendor or operational transition programs.
      • Strong automation mindset with the ability to turn operational processes into repeatable infrastructure and software workflows.
      • Excellent troubleshooting, analytical, and root-cause investigation capabilities.
      • Strong communication and stakeholder-management skills, with the ability to influence technical decisions across teams.
      • Experience mentoring engineers and leading knowledge-transfer initiatives in global or distributed environments.
      • CKA, cloud architecture certifications such as AWS or Azure, or relevant APM/observability certifications are preferred.
      • Benefits:

        • CLT employment with a 40-hour weekly workload.
        • Remote work in Brazil, with the flexibility to work from home when client-site presence is not required.
        • SulAmérica Prestige health insurance, including coverage for eligible legal dependents.
        • SulAmérica dental insurance, including coverage for eligible legal dependents.
        • Prudential life insurance equivalent to 24x salary.
        • MetLife private pension plan with company matching of up to 6%.
        • Flash meal voucher and internet allowance totaling R$1,000 per month.
        • Employee Assistance Program.
        • Wellness program.
        • Base salary plus additional compensation programs, subject to eligibility.
        • Individual performance-based compensation opportunities, where applicable.
        • Equity grant opportunities through the applicable Associate Equity Appreciation Program.
        • Inclusive and diverse working environment with opportunities for professional development and collaboration.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
685,335 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$83k – $204k per year (Estimated) • Remote • Full-Time • 10+ years exp
Python
DevOps
Terraform
AWS
Kubernetes
IAM
Cybersecurity
ISO 27001
SOC 2
Threat Modeling
Web3
Smart Contracts
Staking
Apply
$36k – $88k per year (Estimated) • Remote • Full-Time • 10+ years exp
Python
DevOps
Terraform
AWS
Kubernetes
IAM
Cybersecurity
ISO 27001
SOC 2
Threat Modeling
Web3
Smart Contracts
Staking
Apply
$53k – $122k per year (Estimated) • Remote • Full-Time
Python
JavaScript
PHP
TypeScript
Node JS
Databases
PostgreSQL
ElasticSearch
AI/ML
LangChain
LlamaIndex
Embeddings
Prompt Engineering
Function Calling
AI Agents
CrewAI
LLM
RAG
Semantic Search
Structured Outputs
Semantic Search
LLM Guardrails
Tool Use
Frontend
React.js
Vite
DevOps
Vector
Management
Telegram
WhatsApp
Agile
Apply
$161k – $333k per year (Estimated) • In office • Full-Time
Python
AI/ML
PyTorch
Amazon SageMaker
AWS Trainium
DevOps
AWS
Apply
$54k – $146k per year (Estimated) • In office • Full-Time • Sydney
Python
SQL
DevOps
Azure
AWS
FinOps
Analytics
Power BI
Apply
Legal Counsel aa 5 hours ago
Remote • Full-Time
Apply
$71k – $159k per year (Estimated) • Remote • Full-Time • 2+ years exp
Analytics
Power BI
Looker
Apply
VIP Account Manager 5 hours ago
$58k – $131k per year (Estimated) • Remote • Full-Time • PhD
Apply
Equity • Remote • Full-Time • 8+ years exp
Go
SQL
AI/ML
Claude Code
LLM
DevOps
AWS
Docker
Kubernetes
Management
Agile
Apply
$83k – $204k per year (Estimated) • Remote • Full-Time • 10+ years exp
Python
DevOps
Terraform
AWS
Kubernetes
IAM
Cybersecurity
ISO 27001
SOC 2
Threat Modeling
Web3
Smart Contracts
Staking
Apply
See all jobs
This is one of many
685,335 more open roles from verified company boards, updated every day.