1,341,099open jobs
78,475companies
208,060added this week
Browse all
Salary
≈ $132k – $258k per year (Estimated)
Location
In office (Austin)
Seniority
Staff · 10+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 8, 2026. First seen by Alion on Oct 2, 2026. Virtasant scores B on the Alion truth index.

Overview
Company
Impact
Profile match
We are a global team of cloud experts. We provide an end-to-end set of capabilities to help organizations thrive in the cloud.

Senior/ Staff Platform Engineer

Type: Remote

Coverage: Pacific Hours (8:00 AM - 5:00 PM PST and On-call every 4-5 weeks)

Job Description:

We are looking for a Senior/Staff Platform Engineer to build, operate, and evolve large-scale production infrastructure. This is a hands-on platform and reliability engineering role for someone who has deep experience operating Kubernetes and cloud infrastructure, troubleshooting complex production systems, and building the automation and tooling that keeps those systems reliable.

The work spans Kubernetes, Linux, cloud infrastructure, networking, observability, CI/CD, reliability, and production operations. You will write production code and automation in Go, Python, or Java, but this is not primarily a software development role. We are looking for an engineer who understands the systems underneath the applications and can independently diagnose and solve infrastructure problems across multiple layers.

This is a highly autonomous role. You will work directly with technical stakeholders, own ambiguous infrastructure initiatives from design through production, and be trusted to drive technical decisions and critical issues without requiring constant direction.

Key Responsibilities:

Platform and Kubernetes Engineering:

  • Design, build, operate, and improve production Kubernetes platforms.

  • Own platform-level concerns including cluster architecture, networking, workload isolation, resource management, security, upgrades, scaling, and reliability.

  • Troubleshoot Kubernetes beyond the application layer, including networking/CNI, scheduling, node behaviour, resource constraints, controllers, and cluster-level failures.

  • Operate and improve large-scale, highly available infrastructure across cloud, hybrid, virtualised, and/or bare-metal environments.

  • Diagnose complex issues spanning Kubernetes, containers, Linux, networking, and underlying infrastructure.

  • Optimise platform infrastructure for reliability, performance, scalability, and operational efficiency.

Software Development and Automation:

  • Write, maintain, and improve production tooling and automation using Go, Python, or Java.

  • Build software and automation that improves platform operations, reliability, deployment, troubleshooting, and developer experience.

  • Read, debug, and contribute to existing production codebases.

  • Develop internal services, APIs, integrations, and operational tooling where needed.

  • Automate repetitive operational processes and reduce manual intervention across the platform.

  • Apply sound software engineering practices, including testing, code review, maintainability, and documentation.

Reliability and Production Operations:

  • Own the reliability and operational health of critical production infrastructure.

  • Lead or contribute significantly to incident response for complex platform and infrastructure issues.

  • Investigate root causes and implement durable remediation rather than temporary fixes.

  • Define and improve SLOs, SLIs, alerting, and operational processes.

  • Troubleshoot systems using logs, metrics, traces, profiling tools, and system-level diagnostics.

  • Drive improvements in availability, performance, capacity, resilience, and operational readiness.

  • Contribute to disaster recovery planning, testing, and continuous improvement.

Infrastructure as Code and Delivery:

  • Build and maintain infrastructure as code using Terraform and related automation technologies.

  • Create reusable infrastructure patterns and improve automation as the platform evolves.

  • Build and improve CI/CD and deployment workflows supporting large-scale engineering environments.

  • Balance delivery speed with reliability, security, scalability, and operational requirements.

  • Work across infrastructure provisioning, configuration management, deployment automation, and production operations.

  • Participate in planning and executing production cloud or infrastructure migrations, including dependency analysis, networking, cutover, rollback, and production validation.

Observability:

  • Build and maintain production monitoring, metrics, dashboards, alerting, logging, and distributed tracing.

  • Improve observability so engineers can identify and diagnose failures quickly.

  • Use production telemetry to identify reliability, capacity, and performance problems before they become major incidents.

  • Continuously improve incident detection and reduce time to diagnosis and recovery.

Collaboration and Technical leadership:

  • Work directly with customer and internal engineering teams to understand requirements, investigate problems, and drive technical solutions.

  • Communicate architecture, technical decisions, risks, trade-offs, and progress clearly to technical stakeholders.

  • Own complex infrastructure initiatives from initial problem definition through design, implementation, and production operation.

  • Contribute to architecture discussions, RFCs, design reviews, and technical direction.

  • Mentor other engineers and help improve engineering and operational practices across the team.

  • Operate independently in ambiguous situations and take ownership when immediate technical or management direction is unavailable.

Qualifications:

Education and Experience:

  • 10+ years of professional experience in Platform Engineering, Site Reliability Engineering, Infrastructure Engineering, DevOps, or related fields; 10+ years is preferred for Staff-level candidates.

  • Significant hands-on experience operating complex production infrastructure and distributed systems.

  • Demonstrated experience building and operating production Kubernetes platforms, not only deploying applications onto existing clusters.

  • Production programming experience with Go, Python, or Java.

  • Strong experience with production reliability, incident response, troubleshooting, and operational ownership.

  • Experience independently owning complex technical initiatives from an ambiguous starting point through production.

  • Experience working directly with technical stakeholders or customers and communicating complex technical topics effectively.

  • Degree in Computer Science, Engineering, or a related field, or equivalent practical experience.

Technical Skills:

  • Deep understanding of production Kubernetes infrastructure, including cluster architecture, networking/CNI, NetworkPolicy, scheduling, resource management, nodes, security/RBAC, and cluster behaviour.

  • Strong Linux fundamentals and hands-on production systems troubleshooting.

  • Strong understanding of networking concepts including DNS, routing, load balancing, connectivity, and cloud/Kubernetes networking.

  • Production experience with at least one major cloud platform: AWS, GCP, or Alicloud.

  • Infrastructure as code at scale using Terraform or equivalent tooling.

  • Configuration management and automation experience with technologies such as Ansible, Puppet, or similar.

  • Strong production debugging and root-cause analysis skills across infrastructure and distributed systems.

  • Observability experience using metrics, logs, traces, dashboards, and alerting platforms such as Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent.

  • Experience with CI/CD infrastructure and modern software delivery practices.

  • Docker and container tooling as part of the production lifecycle.

  • Production coding and automation experience in Go, Python, or Java.

  • Understanding of high availability, capacity planning, disaster recovery, and production resilience.

Preferred:

  • Experience planning and executing production cloud or infrastructure migrations, including cutover and rollback strategies.

  • Experience operating large-scale or multi-cluster Kubernetes environments.

  • Experience with hybrid cloud, on-premises, virtualised, or bare-metal infrastructure.

  • Experience building or modifying Kubernetes controllers or operators.

  • Advanced Kubernetes networking, CNI, NetworkPolicy, service mesh, mTLS, or workload identity experience.

  • Multi-cloud infrastructure experience.

  • Experience designing and testing disaster recovery strategies.

  • Experience with large-scale CI/CD or developer infrastructure.

  • Experience with capacity planning and performance engineering.

  • Experience building internal platform tooling or improving developer experience.

  • Experience with security, infrastructure hardening, IAM, or compliance requirements.

  • Previous technical leadership, mentoring, or Staff/Principal-level engineering responsibilities.

  • Experience working directly with external customers or stakeholders in a consulting or service-delivery environment.

Soft Skills:

  • Exceptional written and verbal technical communication.

  • Strong analytical, debugging, and problem-solving ability.

  • High degree of ownership and ability to operate independently.

  • Comfortable making technical decisions and driving work forward in ambiguous environments.

  • Able to communicate effectively with customers, engineers, and technical leadership.

  • Strong technical judgment and ability to explain trade-offs rather than simply implement predefined solutions.

  • Able to lead technically and influence others without requiring formal people-management authority.

  • Comfortable working within a distributed, highly technical team.

Please Note:

  • This is a hands-on platform and reliability engineering role. We are not looking for candidates whose experience has been limited to deploying applications onto Kubernetes, consuming managed cloud services, or provisioning infrastructure without owning its production operation.

  • The ideal candidate has personally built, operated, troubleshot, and improved production infrastructure and can clearly explain what they owned, how the underlying systems worked, and how they approached failures and architectural trade-offs.

  • This role requires significant autonomy and strong technical communication. Engineers should be comfortable working directly with stakeholders and driving complex technical issues without continuous oversight.

  • Production coding experience is required, but it may be in Go, Python, or Java. Go is not a requirement.

  • Cloud migration experience is strongly preferred but is not a knockout requirement.

  • This role is currently open to candidates based in Brazil, Mexico, or Canada.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,341,099 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
Austin
≈ $52k – $113k per year (Estimated) • Hybrid • Full-Time • Hamburg
DevOps
Windows Server
Cybersecurity
Microsoft Entra ID
Active Directory
Apply
In office • Full-Time • 3+ years exp • Bachelor's Degree • Chile
DevOps
SLI/SLO/SLA
Apply
≈ $44k – $112k per year (Estimated) • Hybrid • Full-Time • Austria • Germany
Wolfram
Management
Jira
Apply
DevOps Engineer 21 min ago
≈ $20k – $54k per year (Estimated) • Hybrid • Part-Time • Bachelor's Degree • Cape Town
DevOps
Terraform
Azure DevOps
GitHub Actions
CloudFormation
Azure
CI/CD
AWS
Docker
Kubernetes
Amazon CloudWatch
Apply
$30k – $44k per year • Remote (South Africa) • Full-Time • Cape Town
DevOps
Azure
Management
Google Workspace
Apply
In office • Full-Time • India
Python
Java
C#
Mobile
JUnit
DevOps
CircleCI
GitLab CI
CI/CD
Jenkins
QA
TestNG
Selenium
Appium
Robot Framework
Apply
Hybrid • 8+ years exp • Bachelor's Degree
Python
Java
TypeScript
SQL
C#
C#
.NET
Databases
Databricks
AI/ML
Function Calling
RAG
Semantic Search
OpenAI
GraphRAG
Human-in-the-Loop
Context Engineering
Semantic Search
Knowledge Graph
Tool Use
DevOps
Azure DevOps
Azure
CI/CD
Git
Platform Engineering
Apply
In office • Full-Time • Hyderabad
JavaScript
Java
TypeScript
Java
Spring Boot
Databases
Apache Kafka
AI/ML
Model Context Protocol
Prompt Engineering
AI Agents
AWS Bedrock
RAG
OpenAI
Anthropic
AWS Bedrock AgentCore
AWS Strands Agents
Frontend
React.js
DevOps
AWS
Kubernetes
Apply
Hybrid • 10+ years exp • Bachelor's Degree
Python
SQL
PowerShell
Python
pySpark
Databases
Databricks
AI/ML
Spark
NLP
Time Series Forecasting
Machine Learning
DevOps
Rest API
Azure DevOps
Azure
Kubernetes
Azure AKS
Linux
Management
Agile
Apply
Hybrid • 5+ years exp • Bachelor's Degree
Python
SQL
AI/ML
Machine Learning
DevOps
Kubernetes
Argo Workflows
Analytics
Power BI
ETL/ELT
Microsoft Excel
Management
Power Apps
Microsoft Office
Apply
≈ $119k – $234k per year (Estimated) • Remote (AMER) • Full-Time • Austin
AI/ML
Vertex AI
Fine-tuning
Quantization
AI Agents
AWS Bedrock
LLM
OpenAI
LLM Guardrails
DevOps
Terraform
GCP
Azure
AWS
Kubernetes
FinOps
Apply
≈ $58k – $119k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Austin
DevOps
CI/CD
Apply
≈ $101k – $184k per year (Estimated) • Remote (United States, Canada, PST hours) • Full-Time • 14+ years exp • Austin
SQL
DevOps
GCP
VMWare
CI/CD
AWS
FinOps
Apply
≈ $140k – $255k per year (Estimated) • Remote (United States) • Full-Time • 5+ years exp • Austin
Apply
≈ $96k – $213k per year (Estimated) • In office • Full-Time • 8+ years exp • Austin
SQL
DevOps
CI/CD
Apply
$170k – $205k per year • In office • Full-Time • Bachelor's Degree • Austin
Java
SQL
Databases
Apache Kafka
DevOps
AWS
Kubernetes
Amazon EKS
Linux
Apply
$74k – $83k per year • Remote (United States) • Public Trust • 10+ years exp • Bachelor's Degree • Austin
Python
PowerShell
DevOps
Terraform
Ansible
Azure DevOps
GitHub Actions
Azure
CI/CD
Jenkins
Git
Configuration Management
JFrog Artifactory
GitHub
GitLab
Management
Trello
Confluence
Microsoft Teams
ITIL
Apply
≈ $119k – $223k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • Austin • Durham
Python
Verilog
SystemVerilog
Perl
AI/ML
Quantization
Chips/EDA
UVM
Apply
≈ $51k – $108k per year (Estimated) • Hybrid • Austin
Apply
≈ $58k – $114k per year (Estimated) • In office • Full-Time • 5+ years exp • High School Diploma • Austin
Apply
See all jobs
This is one of many
1,341,099 more open roles from verified company boards, updated every day.