404,711open jobs
14,073companies
78,231added this week
Browse all
Salary
$34k – $85k per year (Estimated)
Location
Remote/Hybrid (Mexico City, Monterrey, Mexico)
Seniority
Senior · 6+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
GE Vernova is an American energy company created in April 2024 when General Electric split into three, taking the power, wind and electrification businesses. It builds and services gas turbines that generate a substantial share of the world's electricity, onshore and offshore wind turbines, grid equipment including transformers and high-voltage transmission, and nuclear technology through GE Hitachi and its BWRX-300 small modular reactor. Headquartered in Cambridge, Massachusetts, the company sits at the centre of two simultaneous demand shocks: the electrification of transport and industry, and the electricity requirements of large-scale artificial intelligence data centres.

Job Description Summary

The Platform System Reliability Engineer is the primary operations engineer and operator of our EKS Kubernetes environment, which serves as the foundation for our global grid software SaaS products. This role focuses on the "middle-mile" of software delivery, ensuring that the underlying compute, networking, and storage layers are secure, hardened, scalable, and resilient to support critical energy infrastructure in the cloud. You will be responsible for the full lifecycle of production clusters, from initial bootstrapping, performance tuning, patching and securing.

Job Description

Roles and Responsibilities

Provision & Infrastructure Hardening

  • Kubernetes Cluster Orchestration: Help design and deploy hardened EKS clusters across multiple AWS regions, ensuring consistent security baselines.

  • Infrastructure as Code (IaC): Build and maintain reusable Terraform and Ansible modules for automated provisioning of cloud infrastructure services including networking services, compute, storage, queue and cache, etc.

  • Security Architecture: Implement "Policy as Code" guardrails and secure network perimeters (ESPs) in alignment with NERC CIP and IEC 62443 standards.

  • Operationalize Cloud Infrastructure: Standardize run books, operating processes required to run critical infrastructure with highest reliability.

Platform Readiness & Scaling

  • Resource Governance: Define and enforce Kubernetes resource quotas, limit ranges, and Pod Priority classes to ensure mission-critical services receive prioritized compute resources.

  • Connectivity & Ingress: Manage the ingress strategy and service mesh architecture to facilitate secure, performant connectivity between distributed micro services.

  • Acceptance Testing: Lead platform-level smoke, load testing and disaster recovery exercises to validate that the infrastructure can meet 99.99% uptime targets.

  • Sizing & Optimization: Partner with application teams to right-size containerized workloads, optimizing for both performance and cloud cost (FinOps).

Operational Excellence & Tier 3 Support

  • L3 Escalation: Act as the highest technical escalation point for complex Kubernetes internals, troubleshooting issues such as failed pods, memory leaks, and network partitions.

  • Incident Response: Lead root cause analysis (RCA) for platform-level outages, implementing systemic fixes to prevent recurring failures.

  • Toil Elimination: Proactively identify and automate repetitive operational tasks-such as cluster upgrades and OS patching-to ensure the team spends at least 50% of their time on engineering improvements.

  • Observability Integration: Institutionalize platform monitoring using Prometheus and Grafana, creating dashboards that surface the "Golden Signals" of cluster health.

Technical Requirements

  • Kubernetes: 5 years of experience operating production-grade Kubernetes clusters at scale.

  • Orchestration & Observability Tools: Expert-level knowledge of multi-cluster management, performance tuning and experience implementing observability tools such as Prometheus/Grafana, Dynatrace, Splunk, Datadog, etc.

  • AWS Infrastructure: Deep hands-on experience with AWS core services (EKS, EC2, ALB, S3, RDS, MSK).

  • Automation Stack: Proficiency in Terraform, Ansible, and Python or Go for infrastructure automation and deployment tools like ArgoCD or Flux.

  • Networking & Security: Strong understanding and hands on experience of cloud networking concepts such as VPCs, routing, load balancing and security configurations such as encryption, certificate management.

  • Fluent in English

Education Qualification

  • Bachelor's Degree in Computer Science, Engineering and Math

Experience

  • Professional Background: 6-8 years in SRE or Platform Engineering roles supporting mission-critical, 24/7 cloud environments.

  • Crisis Management: Proven track record as a structured incident responder who can handle production down/break the glass scenarios in mission critical applications.

Preferred Qualifications

  • Regulated Environments: Practical knowledge of NERC CIP, SOC2, ISO 27001, or IEC 62443 compliance standards in a SaaS context.

  • Certifications: AWS Certified DevOps Engineer - Professional, CKA (CertifiedKubernetes Administrator), or SRE Practitioner Certification.

  • Critical Infrastructure: Experience supporting mission-critical systems in energy, utilities, or other high-stakes industrial sectors.

Business Acumen:

Understand key cross-functional concepts that impact the organization; is aware of business priorities and organizational dynamics

Leadership:

Coach and mentor team members.

Familiar with concepts of costing hardware and software components. Works to assure work is on-time and within budget

Deliver tasks on-time with alignment to architectural goals. Can identify and raise issues, risks and benefits

Participate in change initiatives by implementing new directions and providing appropriate information and feedback

Personal Attributes:

High level of energy and enthusiasm with the ability to thrive in a rapidly changing environment

Demonstrated customer focus - evaluates decisions through the eyes of the customer; builds strong customer relationships; creates processes with customer viewpoint; partners with customers

Change oriented -actively generates process improvements; champions and drives change initiatives; confronts

Ability to work with global teams, act independently and as part of a team

Apply values, policies, procedures and precedent to make timely, routine decisions of limited, clear choice

Open-mindedly to new perspectives or ideas. Consider different or unusual solutions when appropriate

Resolve day-to-day issues related to strategy implementation. Escalate issues that impact the client and/or strategic initiatives

Strong analytical and strong problem solving skills - communicates in a clear and succinct manner and effectively evaluates information/data to make decisions; anticipates obstacles and develops plans to resolve

Additional Information

Relocation Assistance Provided: Yes

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
404,711 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Mexico City
$143k – $169k per year • In office • 3+ years exp • Master's Degree • Los Angeles
C++
Python
SQL
DevOps
Ansible
AWS
Azure
CI/CD
Docker
GCP
Git
Jenkins
Kubernetes
Terraform
Analytics
ETL/ELT
Power BI
Tableau
Apply
$230k – $280k per year • Equity • Remote/Hybrid • Full-Time • 6+ years exp • San Francisco
Python
Databases
DynamoDB
AI/ML
LLM
DevOps
Amazon CloudWatch
Amazon ECS
Amazon EKS
Amazon S3
AWS
AWS Lambda
CI/CD
Kubernetes
Cybersecurity
HIPAA
Apply
$180k – $250k per year • Equity 1–1.5% • Remote • Full-Time • 6+ years exp • Master's Degree • San Francisco
Python
AI/ML
AI Agents
LLM Evaluation
PyTorch
Scale AI
DevOps
AWS
GitHub
Cybersecurity
Zscaler
Management
Dropbox
Marketing
ActiveCampaign
Apply
$220k – $270k per year • Equity 1.2–1.8% • In office • Full-Time • 11+ years exp • San Francisco
Python
TypeScript
AI/ML
AI Agents
Scale AI
DevOps
AWS
GitHub
Kubernetes
Terraform
Cybersecurity
Zscaler
Management
Dropbox
Marketing
ActiveCampaign
Apply
$140k – $154k per year • In office • PhD • Bellevue
MATLAB
Python
SQL
DevOps
AWS
Azure
CI/CD
Analytics
Power BI
Apply
SaaS Cloud Engineer 2 days ago
$23k – $68k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Mexico City • Monterrey
Java
Python
JavaScript
Frontend
Bootstrap
DevOps
Amazon EC2
Amazon EKS
ArgoCD
AWS
CI/CD
CloudFormation
FinOps
GitHub Actions
Jenkins
Kubernetes
Platform Engineering
Progressive Delivery
Rancher
SLI/SLO/SLA
Terraform
Amazon CloudWatch
Amazon S3
GitHub
IAM
Cybersecurity
CVE
FedRAMP
Least Privilege
SOC 2
Apply
$65k – $70k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Schenectady
Apply
$63k – $149k per year (Estimated) • In office • Full-Time • Stafford
MATLAB
MATLAB
Simulink
Apply
$113k – $189k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • United States
IoT
OPC UA
Apply
AI Product Manager 2 days ago
$27k – $65k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Bengaluru
Apply
Assistance Finance 6 hours ago
In office • Full-Time • 2+ years exp • Bachelor's Degree • Mexico City
Apply
$9.6k – $16k per year • Remote • Contractor • PhD • Cape Town • Guayaquil • Santa Cruz • Quezon City • Manila
PHP
PHP
WordPress
Design
Canva
Marketing
HubSpot
Instagram
LinkedIn
YouTube
Apply
$9.6k – $16k per year • Remote • Contractor • 2+ years exp • Monterrey • Guayaquil • Santa Cruz • Cairo • Davao City
PHP
PHP
WordPress
Design
Canva
Marketing
HubSpot
Spotify
YouTube
Apply
$12k – $30k per year • Remote • Contractor • Santo Domingo • Guayaquil • Santa Cruz • Monterrey • Cairo
Design
Canva
Management
Asana
ClickUp
Google Workspace
WhatsApp
Marketing
Instagram
Apply
$12k – $30k per year • Remote • Contractor • Cape Town • Guayaquil • Santa Cruz • Monterrey • Cairo
Analytics
A/B Testing
Management
Google Workspace
Slack
Marketing
Instagram
LinkedIn
Apply
See all jobs
This is one of many
404,711 more open roles from verified company boards, updated every day.