1,011,843open jobs
60,160companies
168,355added this week
Browse all
Salary
≈ $134k – $242k per year (Estimated)
Location
Remote (United States)
Seniority
Staff · 8+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 1, 2026. First seen by Alion on Sep 29, 2026.

Overview
Company
Impact
Profile match
Founded in 2012, EverOps is hired by major SaaS companies to lead the acceleration of their DevOps, SRE, and Network Engineering functions. Our pods of engineers can drop in quickly, make an immediate impact in technically challenging scenarios, a...

Lead Kubernetes Platform Engineer (Remote)

Overview

Some of the world’s most innovative global software and technology companies struggle to find engineering partners capable of stepping into complex environments and immediately driving meaningful outcomes. These teams need more than additional hands-they need senior engineers who can quickly understand an environment, identify the path forward, and execute without constant direction.

Enter EverOps - the premier Embedded Service Provider. We partner directly with customer engineering teams to assess and address mission-critical infrastructure, cloud, and delivery challenges.

The Challenge

EverOps is looking for a Lead Kubernetes Platform Engineer with deep Amazon EKS experience at very large scale to lead a modernization discovery, and the engineering program that follows it, for a high-scale consumer mobile platform.

You’ll be stepping into a multi-cluster production EKS estate where the largest clusters run tens of thousands of pods at peak and compute demand roughly doubles between overnight lows and daytime highs. Upgrading the estate is slow and largely manual, so the team is perpetually behind the Kubernetes release cycle. Compute is one of the largest lines in the business, and the platform backs real-time, safety-critical features where failing to scale at peak is not an option.

This role requires someone who has run Kubernetes at this scale before, can form a clear point of view on an unfamiliar estate quickly, and can turn that into a sequenced, costed plan that engineering leadership can commit to, then lead the team that delivers it.

The Mission

As a Lead Kubernetes Platform Engineer, you will join our U.S.-Based Virtual Operating Center and lead an embedded TechPod as its Pod Leader: a player-coach who sets technical direction, serves as the primary point of contact for the customer’s infrastructure leadership, and stays hands-on in the work.

Your immediate priority is leading a two-month EKS modernization discovery. You’ll baseline the production estate, map Kubernetes version and support status cluster by cluster, analyze blast radius and failure domains, and deliver the upgrade automation design, compute and capacity economics model, and prioritized roadmap needed to commit the program.

As discovery closes, your focus shifts to leading execution of the modernization program across three phases: reducing blast radius through a multi-cluster target architecture, automating the upgrade cycle so the platform stays current without heroics, and right-sizing compute, including a measured, capacity-aware move to Graviton (ARM64).

Success is measured in the customer’s own terms: a lower cost to serve per active user, and a real reduction in the share of platform engineering time consumed by maintenance and keep-the-lights-on work. You will be expected to tie every recommendation back to those outcomes and explain the tradeoffs clearly to engineers and executives alike.

What You’ll Do

  • Estate Discovery: Build a complete baseline of a multi-cluster production EKS estate, covering cluster inventory, topology, workload placement, ownership, per-cluster cost, and Kubernetes version and support status.

  • Blast Radius & Architecture: Analyze failure domains, isolation boundaries, and dependency concentration in very large clusters, and design a multi-cluster target architecture with clear workload placement and tenancy models.

  • Upgrade Automation: Design and implement an upgrade approach (in-place vs. blue/green clusters, Karpenter drift-based node rotation, add-on and API deprecation management) that turns a months-long manual cycle into a repeatable, largely automated process.

  • Compute & Capacity Strategy: Lead instance sizing and workload-fit analysis across instance families, generations, and node sizes, accounting for DaemonSet overhead, bin-packing, network limits, and headroom for rapid scale-up.

  • Graviton Migration: Plan and drive a phased ARM64 migration across Java, Go, PHP, and Python workloads, including multi-architecture builds, native dependency remediation, and per-service rightsizing against latency SLOs.

  • Karpenter Engineering: Configure NodePools, weights, and node overlays so scheduling favors price-performance rather than hourly price, and design safe Spot patterns that protect interruption-sensitive and stream-processing workloads.

  • Cost Engineering: Model compute, support, and commitment economics (Savings Plans, On-Demand, Spot suitability by workload tier, extended support exposure) and quantify savings ranges by lever.

  • Platform Posture: Assess and improve ingress (including migration off Ingress NGINX toward Gateway API), service mesh, CNI, GitOps, and infrastructure-as-code posture across the estate.

  • Operability & Toil Reduction: Quantify maintenance and toil burden using the customer’s own engineering data, set reduction targets, and build the automation that gives time back to platform teams.

  • AI-Native Operations: Define the scope and success criteria for an AI-assisted DevOps agent workstream, and assess platform readiness for AI-accelerated development demand.

  • Technical Leadership: Lead the TechPod as a player-coach, run the working cadence with the customer’s infrastructure owner, and partner with AWS specialists on capacity planning and architecture decisions.

  • Documentation & Readouts: Produce the estate baseline, architecture designs, upgrade plans, roadmap, and executive readout, and present findings and recommendations to engineering leadership.

You Have

  • Experience: 8+ years in DevOps, SRE, Platform, or Infrastructure Engineering, including 4+ years operating production Kubernetes and prior experience in a technical lead, staff, or principal-level role.

  • EKS at Scale: Deep production experience with Amazon EKS at large scale (multiple production clusters, thousands of nodes, or tens of thousands of pods), with a working understanding of where control plane, scheduling, networking, and API limits start to bite.

  • Cluster Upgrades: Hands-on ownership of EKS version upgrades across multiple production clusters, including API deprecations, add-on compatibility, node rotation, and rollback planning, plus a working knowledge of the EKS standard and extended support lifecycle.

  • Karpenter: Advanced production experience with Karpenter (v1+), including NodePools, disruption and consolidation, weighting, and instance-type flexibility.

  • Compute Fundamentals: Strong understanding of EC2 instance families and generations, CPU architecture differences, network performance limits, and how workload bottlenecks shift when the underlying compute changes.

  • ARM64 / Graviton: Experience migrating production workloads to Graviton or other ARM64 platforms, including multi-arch container builds and performance validation.

  • Cloud Cost Management: Experience modeling compute costs and commitments (Savings Plans, Reserved Instances, Spot) and building savings cases grounded in real usage data.

  • Kubernetes Networking: Solid knowledge of the AWS VPC CNI, ingress controllers, Gateway API, service mesh, and load balancing on EKS.

  • Infrastructure as Code: Advanced proficiency with Terraform, plus experience with Helm and GitOps tooling such as Argo CD or Flux.

  • Observability: Comfortable using Datadog, Prometheus, Grafana, or comparable tooling to evaluate cluster health, capacity, and workload performance.

  • Automation: Strong scripting ability using Python, Go, or Bash to automate infrastructure and operational workflows.

  • Discovery & Assessment: Demonstrated ability to enter an unfamiliar environment, surface what isn’t documented, and produce a defensible current-state picture and roadmap in weeks rather than quarters.

  • Communication: Ability to translate technical findings into cost, risk, and capacity terms that engineering leadership and executives can act on.

Extra Awesome

  • Consumer Scale: Experience with high-traffic B2C platforms with sharp daily peaks, or with real-time, safety-critical services where availability at peak is non-negotiable.

  • Multi-Cluster Design: Experience designing cell-based or multi-cluster architectures to reduce blast radius, including fleet management and workload placement across clusters.

  • Price-Performance Benchmarking: Experience running performance-per-dollar comparisons across instance generations and architectures and feeding the results into scheduling policy.

  • Spot at Scale: Experience operating Spot capacity for large production fleets, including interruption handling for Kafka consumers and other stateful or stream-processing workloads.

  • JVM & Runtime Tuning: Experience tuning JVM heap, garbage collection, and Kubernetes requests/limits when moving workloads to new instance generations or architectures.

  • Container-Optimized OS: Experience running Bottlerocket or similar, including node boot-time optimization.

  • AI-Assisted Operations: Experience building AI agents or LLM-driven tooling for infrastructure operations, upgrades, or toil reduction.

  • Consulting: Experience leading assessment or discovery engagements that end in a committed, funded program of work.

  • Certifications: CKA, CKS, AWS Certified Solutions Architect - Professional, AWS Certified DevOps Engineer - Professional, or similar advanced certifications.

Benefits

  • 100% Remote Workplace: We’ve been remote since Day 1!

  • Unlimited Paid Time Off.

  • Equity: Become a true owner of the company.

  • 401K with company contribution and sponsored healthcare.

  • Professional Growth: Access to training and certification programs to accelerate your career.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,011,843 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
In your city
≈ $124k – $247k per year (Estimated) • In office • Top Secret • 10+ years exp • Bachelor's Degree • Amarillo
DevOps
Splunk
Apply
≈ $65k – $128k per year (Estimated) • Equity • Remote (Poland) • Full-Time • Warsaw
AI/ML
LangChain
Semantic Kernel
OpenAI
Anthropic
DevOps
Terraform
GCP
Azure DevOps
Azure
CI/CD
AWS
Kubernetes
Platform Engineering
Azure AKS
FinOps
GitHub
Management
Agile
Apply
$199k – $298k per year • Equity • In office • Full-Time • 10+ years exp • Pleasanton • Atlanta
DevOps
Platform Engineering
Windows
Cybersecurity
Okta
Crowdstrike
ISO 27001
SentinelOne
SOC 2
HIPAA
Zero Trust
Microsoft Entra ID
Apply
$174k – $238k per year • Hybrid • TS/SCI • 8+ years exp • Washington
Python
Go
Java
Java
Apache Tomcat
AI/ML
Machine Learning
DevOps
Terraform
AWS
Docker
Kubernetes
Nginx
Platform Engineering
Amazon CloudWatch
TCP/IP
DNS
VPN
Apache HTTP Server
Cybersecurity
Wireshark
Okta
Tcpdump
FedRAMP
Apply
$174k – $238k per year • In office • TS/SCI • Washington
Python
Go
Databases
MySQL
PostgreSQL
Redis
OpenSearch
AI/ML
Machine Learning
DevOps
Terraform
GCP
Helm
CI/CD
GitOps
ArgoCD
AWS
Kubernetes
Platform Engineering
IAM
DNS
Cybersecurity
Okta
FedRAMP
Apply
$114k – $172k per year • Equity • Hybrid • Full-Time • Master's Degree • Seattle
Python
PowerShell
Bash
DevOps
Terraform
Chef
Apply
≈ $61k – $166k per year (Estimated) • In office • Contractor • Manchester
Python
AI/ML
PyTorch
KServe
DevOps
GCP
CI/CD
Analytics
A/B Testing
Apply
$57k – $67k per year • In office • Full-Time • London • Leeds • Edinburgh
Python
SQL
DevOps
GCP
Azure
AWS
Apply
Business Analyst 2 hours ago
$50k – $77k per year • In office • 2+ years exp • Bachelor's Degree • Rome
Python
SQL
Analytics
Power BI
Microsoft Excel
Apply
≈ $26k – $61k per year (Estimated) • Hybrid • Full-Time • 10+ years exp • Bengaluru
Python
JavaScript
AI/ML
AI Agents
RAG
LLM Guardrails
Multi-Agent Systems
Frontend
React.js
DevOps
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Platform Engineering
GitHub
Management
Agile
ITSM
Apply
≈ $139k – $250k per year (Estimated) • Remote (United States) • Full-Time • 8+ years exp
Python
Go
Databases
Databricks
OpenSearch
AI/ML
Ray
Time Series Forecasting
DevOps
Splunk
Terraform
New Relic
OpenTelemetry
Terragrunt
Datadog
Fluent Bit
Logstash
PagerDuty
Prometheus
AWS
Kubernetes
Grafana
Platform Engineering
Thanos
Mimir
Amazon EKS
Pyroscope
eBPF
Vector
Cortex
Incident Management
SLI/SLO/SLA
Amazon S3
Amazon CloudWatch
Cybersecurity
SOC 2
GDPR
Apply
≈ $102k – $211k per year (Estimated) • Remote (United States) • Full-Time
Databases
Snowflake
DevOps
AWS
Marketing
HubSpot
Apply
See all jobs
This is one of many
1,011,843 more open roles from verified company boards, updated every day.