787,996open jobs
49,912companies
122,545added this week
Browse all
Salary
$135k – $306k per year
Location
In office (Seattle)
Seniority
Principal · 10+ years exp

Confirmed on the employer's own hiring board on Sep 25, 2026. First seen by Alion on Sep 24, 2026. Oracle scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Oracle is an American enterprise technology company founded in 1977 by Larry Ellison, Bob Miner and Ed Oates, and headquartered in Austin, Texas. It built its business on the Oracle Database, still the reference relational engine for large transactional systems, and has expanded into a full applications suite covering finance, human resources, supply chain and customer experience through Fusion Cloud and NetSuite. Its fastest growing segment is Oracle Cloud Infrastructure, which the company has positioned aggressively for AI training and inference workloads through large multi-year capacity contracts.

The Oracle Cloud Infrastructure (OCI) team offers the opportunity to build and operate massive-scale, integrated cloud services in a broadly distributed, multi-tenant cloud environment. OCI builds cloud products for customers who are tackling some of the world's largest technical and business challenges.

Oracle Kubernetes Engine (OKE) is OCI's managed Kubernetes service. OKE enables customers to create, run, scale, secure, and operate Kubernetes clusters on OCI, integrating Kubernetes with OCI compute, networking, storage, identity, observability, security, and automation. The OKE team owns a highly available 24x7 cloud service and is expanding the platform to support larger clusters, higher scale, improved operability, deeper OCI integrations, and increasingly demanding cloud native, AI, and GPU workloads.

We are looking for a senior IC5 software engineer with deep Kubernetes expertise, required cloud infrastructure experience, and a strong distributed systems background. This is a high-impact technical leadership role for an engineer who can define architecture, drive cross-team execution, solve ambiguous production and platform problems, and deliver durable systems that improve both customer experience and operational excellence.

You will work on core OKE platform capabilities including cluster lifecycle management, orchestration, scalability, reliability, performance, automation, observability, security, and integration with OCI infrastructure services. The ideal candidate has hands-on experience designing, building, operating, or deeply debugging production cloud services, infrastructure platforms, or Kubernetes-based systems at meaningful scale.

This role requires advanced Kubernetes experience, including Kubernetes control plane behavior, controllers and operators, scheduling, autoscaling, networking, storage, service discovery, container runtimes, node lifecycle, Kubernetes APIs, and etcd. Experience with Kubernetes networking and storage technologies such as CNI, Cilium, Calico, Flannel, other container networking implementations, CSI drivers, and cloud provider integrations is highly relevant.

OKE is also expanding to support demanding AI and accelerated computing use cases. Experience with AI/ML infrastructure, multi-node GPU clusters, accelerated compute, model training or inference platforms, GPU scheduling, device plugins, Karpenter, cluster autoscaling, CUDA, NCCL, RoCE, InfiniBand, RDMA, SmartNIC/DPU offload, or high-performance AI/HPC networking is a significant plus.

This role also requires an engineer who is ready to use modern agentic engineering practices responsibly. We expect senior engineers to apply AI-assisted and agentic workflows to accelerate design exploration, implementation, testing, debugging, documentation, operational analysis, and developer productivity while maintaining strong ownership, security judgment, code quality, and production accountability.

As a member of the software engineering division, you will take an active role in defining and evolving standard practices and procedures. You will define specifications for significant new projects and specify, design, develop, troubleshoot, and debug software for OCI's managed Kubernetes service.

Responsibilities include:

  • Provide technical leadership for major OKE platform initiatives from architecture through implementation, launch, and production operation.
  • Design and build distributed systems that create, update, scale, repair, and operate Kubernetes clusters across OCI regions.
  • Improve OKE reliability, scalability, performance, upgrade safety, lifecycle management, observability, automation, and operational tooling.
  • Work deeply with Kubernetes technologies, including control plane components, controllers/operators, scheduling, autoscaling, Kubernetes APIs, container runtimes, node behavior, and etcd.
  • Design, debug, and improve Kubernetes networking and storage integrations, including CNI-based networking, Cilium, Calico, Flannel, other container networking implementations, CSI drivers, and OCI infrastructure integrations.
  • Build automation for cluster validation, health checks, readiness testing, failure detection, remote recovery, and reduction of post-deployment operational issues.
  • Lead technical design reviews, code reviews, incident reviews, and production readiness reviews for complex service changes.
  • Debug difficult production issues across service boundaries, including Kubernetes, Linux, networking, compute, storage, identity, telemetry, and OCI infrastructure dependencies.
  • Apply performance engineering practices including profiling, tracing, latency analysis, throughput optimization, and production diagnostics across distributed systems.
  • Build automation that reduces manual operations, improves fleet health, accelerates diagnosis, and raises the quality bar for OKE engineering.
  • Partner with OCI service teams to deliver end-to-end platform capabilities regardless of organizational boundaries.
  • Apply AI-assisted and agentic engineering workflows to improve engineering velocity, test coverage, debugging, operational analysis, and documentation while ensuring correctness, security, and maintainability.
  • Mentor engineers, influence technical direction, and help establish patterns that scale across the OKE organization.
  • Participate in operating a 24x7 cloud service and use customer feedback, production data, and operational experience to prioritize improvements.

Required qualifications:

  • 10+ years of software engineering experience, or equivalent experience building and operating production software systems.
  • Hands-on cloud infrastructure experience is required, ideally designing, building, operating, or debugging production services or platforms on OCI, AWS, Azure, GCP, or a large-scale private cloud.
  • Strong hands-on Kubernetes expertise is required, including Kubernetes architecture, APIs, control plane behavior, controllers/operators, scheduling, autoscaling, networking, storage, nodes, cluster lifecycle management, or production cluster operations.
  • Advanced Kubernetes knowledge, including CNI, CSI, etcd, service discovery, container runtimes, node lifecycle, and Kubernetes failure modes.
  • Experience with Kubernetes networking technologies such as Cilium, Calico, Flannel, or other CNI implementations.
  • Experience with Kubernetes storage integrations, including CSI drivers or cloud storage integrations.
  • Strong distributed systems fundamentals, including availability, failure handling, performance, scalability, and operational tradeoffs.
  • Experience building highly available infrastructure services, platform services, or cloud native systems used in production.
  • Strong development experience in both Go/Golang and Java is required.
  • Strong Linux, networking, debugging, and production operations skills.
  • Demonstrated ability to lead ambiguous technical projects, influence across teams, and deliver through other engineers without relying on formal authority.
  • Strong communication skills, ownership, judgment, and ability to make pragmatic tradeoffs in production systems.

Preferred qualifications:

  • Experience with AI/ML infrastructure, GPU workloads, multi-node GPU clusters, accelerated compute, model training or inference platforms, GPU scheduling, device plugins, Karpenter, cluster autoscaling, CUDA, NCCL, high-performance networking, or distributed training systems.
  • Experience with eBPF-based networking, Kubernetes network policy, service mesh, ingress, load balancing, overlays/underlays, BGP, VXLAN, SmartNIC/DPU offload, RoCE, InfiniBand, RDMA, or multi-cluster networking.
  • Experience with infrastructure as code and cloud provisioning tools such as Terraform, Packer, cloud-init, IAM, VCN/VPC networking, VPN, FastConnect/Direct Connect, or equivalent cloud primitives.
  • Experience building developer productivity, operational automation, or responsible AI-assisted and agentic engineering workflows.
  • Experience with observability systems, incident response, safe deployment practices, canary analysis, rollback strategies, service health automation, and large fleet operations.
  • Open-source or upstream contribution experience in Kubernetes, cloud native infrastructure, observability, networking, or related systems.
Disclaimer:

Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.

Range and benefit information provided in this posting are specific to the stated locations only

US: Hiring Range in USD from: $135,200 to $306,400 per annum. May be eligible for bonus, equity, and compensation deferral.

Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.

Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following:

1. Medical, dental, and vision insurance, including expert medical opinion

2. Short term disability and long term disability

3. Life insurance and AD&D

4. Supplemental life insurance (Employee/Spouse/Child)

5. Health care and dependent care Flexible Spending Accounts

6. Pre-tax commuter and parking benefits

7. 401(k) Savings and Investment Plan with company match

8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.

9. 11 paid holidays

10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.

11. Paid parental leave

12. Adoption assistance

13. Employee Stock Purchase Plan

14. Financial planning and group legal

15. Voluntary benefits including auto, homeowner and pet insurance

The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.

Career Level - IC5

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
787,996 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
Seattle
≈ $140k – $261k per year (Estimated) • Remote (United States) • Full-Time • 12+ years exp • United States
JavaScript
Databases
PostgreSQL
Frontend
React.js
DevOps
AWS
Kubernetes
Apply
$150k – $262k per year • Equity • Hybrid • Full-Time • 6+ years exp • Waltham
Python
JavaScript
AI/ML
Prompt Engineering
AI Agents
Management
ServiceNow
Apply
$221k – $387k per year • Equity • Hybrid • Full-Time • Bachelor's Degree • New York
Python
AI/ML
Claude
Function Calling
AI Agents
LLM Guardrails
Agentic Workflows
Tool Use
Machine Learning
DevOps
GCP
Azure
AWS
Docker
Kubernetes
Management
ServiceNow
Apply
$150k – $262k per year • Equity • Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Kirkland
DevOps
Terraform
GCP
Azure
CI/CD
GitOps
AWS
Kubernetes
Platform Engineering
Service Mesh
Amazon EKS
Google GKE
Azure AKS
IAM
Management
ServiceNow
Apply
$161k – $250k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • Salt Lake City
Python
JavaScript
Java
TypeScript
C#
AI/ML
AWS Bedrock
LLM
RAG
Hybrid Search
Amazon SageMaker
LLM Guardrails
Frontend
Angular
DevOps
Rest API
CI/CD
AWS
Cybersecurity
HIPAA
Apply
≈ $79k – $211k per year (Estimated) • In office • 3+ years exp • Master's Degree • Singapore
AI/ML
Stable Diffusion
DeepSpeed
CUDA Toolkit
Diffusion Models
Transformers
PyTorch
CUDA
Pre-training
Megatron-LM
NCCL
cuDNN
Apply
Engineering Manager 10 hours ago
≈ $39k – $95k per year (Estimated) • In office • Full-Time • Noida
JavaScript
TypeScript
Node JS
Databases
Redis
ClickHouse
Memcached
Apache Kafka
DevOps
Rest API
gRPC
GCP
DigitalOcean
GitHub Actions
Jenkins
Docker
Kubernetes
Grafana
Apply
≈ $109k – $245k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Denver • Phoenix
AI/ML
NLP
LLM
DevOps
GCP
Azure
AWS
Cybersecurity
Darktrace
SIEM
Apply
≈ $88k – $173k per year (Estimated) • Hybrid • Full-Time • Bachelor's Degree • London
DevOps
GCP
Azure
AWS
Cybersecurity
ISO 27001
MITRE ATT&CK
Apply
≈ $20k – $49k per year (Estimated) • Remote (India) • Full-Time • 8+ years exp • Bachelor's Degree • Hyderabad
Python
AI/ML
Fine-tuning
Prompt Engineering
AI Agents
LLM
LLMOps
Multi-Agent Systems
DevOps
GCP
Azure
AWS
Apply
$87k – $187k per year • Equity • In office • 5+ years exp • Bachelor's Degree • Austin
Python
JavaScript
SQL
C++
Perl
DevOps
Linux
Unix
Apply
≈ $119k – $221k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Reston
DevOps
Azure
AWS
Incident Management
Apply
$93k – $210k per year • Equity • In office • Austin
Java
C#
C++
Cybersecurity
PKI
Apply
$93k – $210k per year • Equity • In office • 3+ years exp • Bachelor's Degree • Nashville
Python
Go
Java
C#
C++
DevOps
CI/CD
GitOps
Kubernetes
Platform Engineering
Apply
Software Developer 5 11 days ago
$135k – $306k per year • Equity • In office • 8+ years exp • Bachelor's Degree • Nashville
Python
Java
C++
Apply
$144k – $194k per year • Equity • In office • Full-Time • 3+ years exp • Bachelor's Degree • Seattle
Java
C#
C++
Perl
DevOps
AWS
Amazon EC2
Amazon S3
Apply
≈ $144k – $265k per year (Estimated) • Equity • In office • Full-Time • 5+ years exp • Bachelor's Degree • Seattle
Java
C#
C++
Perl
Databases
Apache Iceberg
DynamoDB
AI/ML
Spark
Machine Learning
DevOps
AWS
Amazon Kinesis
Apply
$144k – $194k per year • Equity • In office • Full-Time • 3+ years exp • Bachelor's Degree • Seattle
Java
C#
C++
Perl
Management
Agile
Apply
Account Executive 10 hours ago
$75k – $150k per year • Remote (United States) • Full-Time • 1+ year exp • Seattle
Cybersecurity
DLP
Design
Webflow
Management
UiPath
Stripe
Marketing
LinkedIn
Apply
$140k – $200k per year • Remote (United States, Indonesia, Malta) • Full-Time • 1+ year exp • Seattle
Apply
See all jobs
This is one of many
787,996 more open roles from verified company boards, updated every day.