998,188open jobs
59,534companies
165,604added this week
Browse all
Salary
≈ $139k – $250k per year (Estimated)
Location
Remote (United States)
Seniority
Staff · 8+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 1, 2026. First seen by Alion on Sep 29, 2026.

Overview
Company
Impact
Profile match
Founded in 2012, EverOps is hired by major SaaS companies to lead the acceleration of their DevOps, SRE, and Network Engineering functions. Our pods of engineers can drop in quickly, make an immediate impact in technically challenging scenarios, a...

Lead Observability Engineer (Remote)

Overview

Some of the world’s most innovative global software and technology companies struggle to find engineering partners capable of stepping into complex environments and immediately driving meaningful outcomes. These teams need more than additional hands-they need senior engineers who can quickly understand an environment, identify the path forward, and execute without constant direction.

Enter EverOps - the premier Embedded Service Provider. We partner directly with customer engineering teams to assess and address mission-critical infrastructure, cloud, and delivery challenges.

The Challenge

EverOps is looking for a Lead Observability Engineer with deep hands-on experience across metrics, logs, and traces at very large scale to lead an observability maturity assessment, and the platform consolidation that follows it, for a high-scale consumer mobile platform.

The current estate spans multiple commercial observability vendors alongside self-managed Prometheus, Thanos, Grafana, Vector, and ELK. It carries tens of millions of active time series, tens of thousands of scrape targets, well over a thousand dashboards, and more than ten thousand alert definitions. Much of the log and trace data is sampled or dropped for cost reasons, ownership tagging is sparse, and the self-managed components carry an operational load that crowds out improvement.

The direction under evaluation is consolidation onto AWS-native observability services built on OpenTelemetry. This role requires someone who knows these systems well enough to price that move honestly, prove or disprove performance parity, and then lead the migration if the answer is go.

The Mission

As a Lead Observability Engineer, you will join our U.S.-Based Virtual Operating Center and lead an embedded TechPod as its Pod Leader: a player-coach who sets technical direction, serves as the primary point of contact for the customer’s engineering leadership, and stays hands-on in the work.

Your immediate priority is leading a two-month Observability Maturity Assessment. You’ll validate telemetry volumes, retention, sampling, and cardinality; build a total cost of ownership model; design an AWS-native target architecture; and deliver a migration plan and commercial recommendation that leadership can make a go/no-go decision on.

As discovery closes, your focus shifts to leading the platform migration: standing up the target collection pipeline and storage tiers, porting log transforms to OpenTelemetry, rebuilding dashboards and alerts, and establishing policy-driven retention and ownership tagging that make observability both less expensive and more useful.

The customer’s observability team is capable but stretched thin. You will be expected to add capacity rather than consume it: pull the data yourself, ask targeted questions, keep recommendations unbiased, and leave the team with fewer systems to run and better telemetry than they have today.

What You’ll Do

  • Estate Assessment: Build a complete picture of the observability estate across vendors, agents, collectors, query surfaces, data volumes, and operating model, grounded in live measurement rather than questionnaires alone.

  • Telemetry Analysis: Validate metric cardinality, active series, scrape target health, log volumes, trace sampling, and retention, and identify where fidelity is being lost today and why.

  • Cost Modeling: Build a total cost of ownership comparison between the current state and two to three costed target states, covering vendor spend, self-managed infrastructure, data transfer and egress, and AWS Pricing Calculator estimates.

  • Target Architecture: Design the AWS-native target across collection (ADOT / OpenTelemetry Collector), pipeline (Amazon Data Firehose), storage tiering (CloudWatch, S3, Amazon Managed Service for Prometheus), query surfaces (Amazon Managed Grafana, CloudWatch, Athena, OpenSearch), and alerting.

  • Retention & Data Classification: Partner with Security, Legal, and Engineering to classify telemetry data and design policy-driven retention, including long-term compliance archives.

  • Capability & Gap Analysis: Map current capabilities to AWS-native equivalents and make clear recommendations where no equivalent exists, such as continuous profiling.

  • Performance Parity: Define and run tests that show whether the target state holds query performance, alert latency, and data fidelity against success criteria agreed with engineering leadership.

  • Migration Planning & Execution: Size and sequence the migration, then lead it, including pipeline cutover, porting log transforms to OpenTelemetry, rebuilding dashboards, and deduplicating and rebuilding alerts.

  • OpenTelemetry Standards: Define the OpenTelemetry conventions (semantic conventions, resource attributes, collector topology, sampling strategy) that application teams will adopt as instrumentation moves over.

  • Ownership & Cost Attribution: Rebuild service ownership tagging so telemetry cost can be attributed back to the teams generating it.

  • Commercial Analysis: Reconcile platform spend, analyze licensing and commit structures, and inform vendor renewal strategy with a clear, unbiased recommendation.

  • Technical Leadership: Lead the TechPod as a player-coach, run the working cadence with the customer’s observability leadership, and keep the engagement light-touch on a busy internal team.

  • Documentation & Readouts: Produce the assessment report, target architecture, cost model, migration plan, and executive summary, and present findings and tradeoffs to engineering leadership.

You Have

  • Experience: 8+ years in SRE, DevOps, Observability, or Platform Engineering, including 4+ years owning production observability platforms and prior experience in a technical lead, staff, or principal-level role.

  • Observability at Scale: Deep experience running metrics, logging, and tracing platforms at large scale (millions of active series, multiple terabytes of logs per day) and making them cheaper and more reliable over time.

  • Prometheus Ecosystem: Advanced production experience with Prometheus and a long-term storage layer such as Thanos, Cortex, or Mimir, including PromQL, recording rules, remote write, and cardinality management.

  • Commercial Platforms: Hands-on experience with Datadog or a comparable commercial platform, including how its pricing is built across hosts, custom metrics, log indexing, and APM.

  • AWS Observability: Production experience with CloudWatch, CloudWatch Logs, Container Insights, X-Ray or Application Signals, Amazon Managed Service for Prometheus, and Amazon Managed Grafana.

  • OpenTelemetry: Production experience with the OpenTelemetry Collector (or ADOT), including pipeline design, processors, tail and head sampling, and multi-backend export.

  • Log Pipelines: Experience designing and operating log pipelines with Vector, Fluent Bit, Logstash, or Firehose, including transforms, routing, and tiered storage in S3.

  • Kubernetes: Strong production experience with EKS or Kubernetes, including DaemonSet agent sizing, kube-state-metrics, and monitoring very large clusters.

  • Infrastructure as Code: Advanced proficiency with Terraform; experience with Terragrunt or Atmos is a plus.

  • Alerting & Incident Response: Experience designing SLO-based alerting, reducing alert sprawl, and integrating with incident management tooling such as PagerDuty.

  • Cost Modeling: Ability to build defensible TCO models from usage data and pricing, and to explain the assumptions behind every number.

  • Automation: Strong scripting ability using Python, Go, or Bash to pull usage data, analyze telemetry, and automate migration work.

  • Discovery & Assessment: Demonstrated ability to enter an unfamiliar environment, measure it directly, and produce a defensible current-state picture and recommendation in weeks rather than quarters.

  • Communication: Ability to explain technical, cost, and compliance tradeoffs to engineers, Security and Legal stakeholders, and executive leadership.

Extra Awesome

  • Vendor Migration: Experience migrating off Datadog, Splunk, New Relic, or similar platforms onto AWS-native or open-source observability stacks.

  • Dashboards & Alerts as Code: Experience automating large-scale dashboard and alert migrations using Grafana provisioning, Grafonnet, Terraform providers, or conversion tooling.

  • Grafana Ecosystem: Experience with Grafana Alloy, Beyla, or other eBPF-based instrumentation.

  • Continuous Profiling: Experience with Pyroscope, Parca, or comparable profiling tools.

  • Data Governance: Experience designing telemetry retention and data handling to meet SOC 2, GDPR, privacy, or similar compliance requirements.

  • Analytics Platforms: Familiarity with Athena, OpenSearch, Databricks, or similar platforms used for log analytics and long-term telemetry queries.

  • Consumer Scale: Experience with high-traffic B2C platforms where telemetry volume tracks tens of millions of users.

  • Consulting: Experience leading assessment or discovery engagements that end in a committed, funded program of work.

  • Certifications: Prometheus Certified Associate, OpenTelemetry Certified Associate, AWS Certified DevOps Engineer - Professional, AWS Certified Solutions Architect - Professional, CKA, or similar.

Benefits

  • 100% Remote Workplace: We’ve been remote since Day 1!

  • Unlimited Paid Time Off.

  • Equity: Become a true owner of the company.

  • 401K with company contribution and sponsored healthcare.

  • Professional Growth: Access to training and certification programs to accelerate your career.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
998,188 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
In your city
≈ $119k – $245k per year (Estimated) • Hybrid • Full-Time • 15+ years exp • Bachelor's Degree • Newberg
Apply
$107k – $230k per year • Remote (United States)
DevOps
Terraform
Ansible
OpenShift
Rancher
CI/CD
GitOps
Kubernetes
Linux
Apply
Team Lead DevOps 3 hours ago
≈ $22k – $47k per year (Estimated) • Remote (likely EAEU) • 5+ years exp • Saint Petersburg
Databases
PostgreSQL
Redis
ClickHouse
Apache Kafka
DevOps
Ansible
Loki
VMWare
Prometheus
GitLab CI
CI/CD
Kubernetes
Grafana
Sealed Secrets
SLI/SLO/SLA
Linux
Apply
$185k – $200k per year • Equity • Remote (United States) • 5+ years exp • Bachelor's Degree
Python
Java
AI/ML
Copilot
Claude
DevOps
Terraform
Ansible
GCP
Chef
CI/CD
Git
Docker
Kubernetes
Configuration Management
Management
n8n
Apply
$212k – $321k per year • Equity • Hybrid • PhD • Boston
Verilog
DevOps
HPC
Apply
Engineer II - DevOps 2 hours ago
≈ $11k – $29k per year (Estimated) • Hybrid • Full-Time • Pune
Python
PowerShell
AI/ML
Copilot
AI Agents
DevOps
Terraform
Ansible
Helm
GitHub Actions
Terragrunt
Datadog
Prometheus
GitLab CI
CI/CD
GitOps
ArgoCD
Jenkins
AWS
Kubernetes
Grafana
Platform Engineering
Cybersecurity
Sumo Logic
Apply
In office • 8+ years exp • Bengaluru
Python
Go
Java
TypeScript
Python
FastAPI
Asyncio
Celery
Pydantic
Mypy
Databases
PostgreSQL
Redis
RabbitMQ
Apache Kafka
AI/ML
LangGraph
AutoGen
LangChain
LlamaIndex
AI Agents
Instructor
Langfuse
LLM
RAG
Human-in-the-Loop
Structured Outputs
LLM Evaluation
LLM Guardrails
Multi-Agent Systems
Tool Use
DevOps
OpenTelemetry
CI/CD
Apply
≈ $13k – $32k per year (Estimated) • Equity • Hybrid • Full-Time • 1+ year exp • Bangkok
PowerShell
C#
Bash
COBOL
COBOL
IBM MQ
Databases
MS SQL
DevOps
Terraform
Puppet
Ansible
Helm
GitHub Actions
Chef
Datadog
Azure
CI/CD
Git
Kubernetes
Configuration Management
Linux
Windows
Management
Agile
QA
Postman
SoapUI
Apply
≈ $22k – $45k per year (Estimated) • In office • 3+ years exp • Moscow
Go
SQL
Databases
PostgreSQL
ClickHouse
DevOps
CI/CD
Docker
Kubernetes
Linux
Apply
Full Stack Developer 2 hours ago
In office • Full-Time • 4+ years exp • Bachelor's Degree • Bengaluru
Python
JavaScript
TypeScript
SQL
Python
Flask
Django
Databases
PostgreSQL
DynamoDB
Amazon Aurora
Frontend
Zustand
Redux
React.js
Mobile
State Management
DevOps
Terraform
GitHub Actions
AWS CDK
CI/CD
Jenkins
Git
AWS
Docker
AWS Lambda
Amazon EC2
GitHub
Amazon S3
API Gateway
Apply
≈ $134k – $242k per year (Estimated) • Remote (United States) • Full-Time • 8+ years exp
Python
Java
PHP
Databases
Apache Kafka
AI/ML
AI Agents
LLM
DevOps
Terraform
Helm
Datadog
Prometheus
GitOps
ArgoCD
AWS
Kubernetes
Grafana
Platform Engineering
Service Mesh
Karpenter
Gateway API
ingress-nginx
Amazon EKS
Amazon EC2
Apply
≈ $102k – $211k per year (Estimated) • Remote (United States) • Full-Time
Databases
Snowflake
DevOps
AWS
Marketing
HubSpot
Apply
See all jobs
This is one of many
998,188 more open roles from verified company boards, updated every day.