368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$146k – $306k per year
Location
In office (Nashville)
Seniority
Principal
Overview
Company
Impact
Profile match
Oracle Corporation is an American multinational computer technology corporation headquartered in Austin, Texas. Founded in 1977 by Larry Ellison, Bob Miner, and Ed Oates, Oracle is one of the world's largest enterprise software and cloud computing infrastructure providers.

Mentors teams and leads the architecture of highly scalable, interdependent distributed systems. Identifies and removes performance/scalability bottlenecks for hyper-scale workloads; defines scalability requirements with stakeholders; and designs elastic, high-impact systems while advancing innovation in data plane platforms. Engineers and oversees fault-tolerant, in-service-upgradable designs; optimizes resilience mechanisms (load-shedding, throttling, rate-limiting); and sets SLO-aligned durability and availability standards across dependent services. Establishes KPIs and advanced telemetry; applies formal verification for complex features; and develops robust replication/synchronization strategies. Advises and leads resolution of complex production issues, sets operational readiness and SOP standards, and directs incident response and RCAs. Architects advanced security controls, drives remediation and compliance, and delivers enterprise-level automation (IaC) and change strategies enabling safe, automated patching, updates, and rollbacks.

Oracle Cloud Infrastructure (OCI) is building the next generation of cloud infrastructure for customers around the world. Our Network Automation team develops the software systems that enable OCI to design, build, scale, and operate its global network.

We are seeking a Lead Principal Software Engineer to help build and operate software for network control- and management-plane services. In this role, you will work on highly available, distributed systems that manage and automate operations across hundreds of thousands of network devices.

You will be a hands-on technical leader who can independently drive complex projects, influence architecture and operational practices, and partner across engineering, operations, quality assurance, and vendor teams. You will help improve the reliability, scalability, performance, and operability of critical OCI networking services.

Key Responsibilities

System Design & Architecture - System Scalability:

-Mentor the team in the architecture and design of highly scalable, interdependent distributed systems, ensuring horizontal and vertical scalability and overall performance, including leveraging distributed state management tools.

-Lead the identification of performance and scalability bottlenecks and recommend solutions to optimize code and/or systems for large-scale data processing and high-throughput requirements to improve performance for hyper-scale systems.

-Lead collaboration with stakeholders to define system scalability requirements, ensuring the defined requirements meet customer expectations.

-Leverage deep expertise to design high-impact, interdependent systems to scale with elasticity (e.g., effectively scaling both up and down).

-Drive innovation in the use of data plane platforms.

-Evaluate whether systems are meeting nonfunctional scalability requirements, and proactively anticipate growing business needs within the business unit.

System Design & Architecture - System Reliability Design:

-Design and oversee the implementation of fault-tolerant, interdependent systems capable of withstanding in-service updates by implementing sophisticated redundancy, replication, and automatic failover capabilities.

-Lead the design and implementation of systems that effectively handle service disruptions (e.g., network partitions) by prioritizing consistency, availability, or partition tolerance.

-Guide the optimization of advanced mechanisms to handle network unreliability, including load-shedding, throttling, and rate-limiting.

-Design interdependent systems that are durable and adhere to service level objectives (SLOs), driving standards for availability and durability of other computing services within the organization

System Design & Architecture - System Reliability Performance:

-Define key performance indicators (KPIs) and telemetry to identify risks, gaps, or cyclical dependencies in running, interdependent systems.

-Drive the creation and customization of highly complex dashboards, telemetry systems, and alerting mechanisms, proactively ensuring system health and reliability.

System Design & Architecture - Correctness / Availability:

-Maintain expertise in industry standards for verifying correctness and apply existing techniques to interdependent systems.

-Formally verify complex features (e.g., via TLA+) to ensure system design correctness for various interdependent systems.

-Develop advanced strategies for data replication and synchronization, ensuring robust data integrity and availability

Operational Troubleshooting & Incident Management:

-Advise on efforts to diagnose, debug, and resolve complex issues in active, interdependent systems to support ongoing operation.

-Develop and implement comprehensive strategies to prevent interruptions, ensuring no maintenance windows are required for customers and users when resolving issues.

-Maintain expertise in dependencies, dependents, and owned systems to drive effective troubleshooting and performance.

-Set standards for operational readiness and standard operating procedures within the department, and hold third-party partners accountable for meeting those standards.

-Oversee operational support rotations, providing expert guidance in incident response and leading root cause investigations to prevent future occurrences.

Compliance & Security:

-Architect advanced security measures to protect data and applications in multi-tenant environments, and lead initiatives to enhance data and application protection.

-Guide the execution of comprehensive remediation plans to address identified security vulnerabilities.

-Ensure cloud infrastructure is in compliance with industry standards and regulations, and guide documentation efforts across projects.

Automation & Change Management:

-Develop enterprise-level automation tools and strategies (e.g., Infrastructure as Code (IaC)) and oversee their implementation.

-Drive alignment of change management plans and organizational initiatives for patching, updating, and rolling back applications, and design interdependent systems to allow for automation of these processes.

Core Responsibilities

Planning & Execution:

-Manages and provides direction on timelines, deliverables, and budgets when applicable for critical high-impact projects or initiatives that impact the line of business, ensuring timely completion and adherence to requirements. Anticipates and plans for shifts in resources or timelines based on changing business priorities, ensuring optimal outcomes.

Collaboration & Partnership:

-Influences cross-functional leaders and external stakeholders to gain alignment on strategic objectives. Fosters partnerships with key business leaders, stakeholders, and/or customers, identifying opportunities for expanding partnerships and promoting long-term organizational success. Champions transparency and inclusivity by actively seeking, listening to, and incorporating diverse perspectives.

Problem Solving:

-Leads specialized, advanced problem-solving efforts, serving as an escalation point for complex issues. Guides others to leverage innovative data-driven techniques to address ambiguous or novel issues, identify root causes, and drives the implementation of solutions that prevent future issues.

Continuous Learning:

-Leverages deep industry knowledge and expertise to serve as a thought leader within the organization. Contributes to the advancement of the field or industry through thought leadership (e.g., conference presentations, white papers, research contributions). Maintains and evolves expertise in relevant areas by proactively monitoring emerging trends, technologies, and industry standards, ensuring the organization remains current with best practices. Champions continuous learning and knowledge sharing, promoting professional development across teams. Applies new knowledge to drive advancement and mentors others to do the same.

Continuous Improvement:

-Develops innovative solutions and drives the implementation of ideas that increase the efficiency and effectiveness of processes, protocols, and workflows across the organization. Evaluates effectiveness of updated approaches and methods for continued improvement to enhance efficiencies and ensure changes align with organizational goals. Designs and develops metrics to measure success of improvement initiatives.

Performance and Development:

-Serves as a subject matter expert regarding talent needs and organizational talent strategy. Imparts leadership and expert knowledge throughout the talent development pipeline including candidate interviews, candidate assessment, and hiring decisions, ensuring alignment with organizational talent strategy.

Disclaimer:

Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.

Range and benefit information provided in this posting are specific to the stated locations only

US: Hiring Range in USD from: $146,300 to $306,400 per annum. May be eligible for bonus, equity, and compensation deferral.

Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.

Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following:

1. Medical, dental, and vision insurance, including expert medical opinion

2. Short term disability and long term disability

3. Life insurance and AD&D

4. Supplemental life insurance (Employee/Spouse/Child)

5. Health care and dependent care Flexible Spending Accounts

6. Pre-tax commuter and parking benefits

7. 401(k) Savings and Investment Plan with company match

8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.

9. 11 paid holidays

10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.

11. Paid parental leave

12. Adoption assistance

13. Employee Stock Purchase Plan

14. Financial planning and group legal

15. Voluntary benefits including auto, homeowner and pet insurance

The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.

Career Level - IC5

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Nashville
$252k – $308k per year • Remote/Hybrid • Full-Time • 7+ years exp • Mountain View
Python
Node JS
JavaScript
Node JS
Commander.js
Databases
Apache Kafka
DynamoDB
AI/ML
AI Agents
ChatGPT
Claude
Claude Code
Copilot
Cursor
DevOps
Amazon EKS
AWS
Datadog
FinOps
Incident Management
Kubernetes
OpenTelemetry
SLI/SLO/SLA
Terraform
Amazon CloudWatch
Cybersecurity
SOC 2
Management
Slack
Apply
$145k – $295k per year (Estimated) • Equity • Remote • Internship • 8+ years exp • San Francisco
Go
Databases
Apache Kafka
ClickHouse
Google BigQuery
AI/ML
Flink
Recommender Systems
DevOps
Incident Management
Kubernetes
Apply
$258k – $386k per year • Equity • Remote/Hybrid • Full-Time • 8+ years exp • San Francisco
Go
Java
Python
Rust
Databases
DynamoDB
MySQL
PostgreSQL
Redis
DevOps
AWS
Incident Management
Kubernetes
Apply
In office • 3+ years exp • Bachelor's Degree • Phoenix
PowerShell
Python
DevOps
SLI/SLO/SLA
Cybersecurity
Okta
Tanium
Management
Google Workspace
Slack
ServiceNow
Apply
$22k – $45k per year (Estimated) • In office • Moscow
AI/ML
LLM
DevOps
SLI/SLO/SLA
Apply
$115k – $235k per year • Equity • In office • Bachelor's Degree • Nashville
Go
Java
Python
AI/ML
AutoGen
Claude
Claude Code
Copilot
CrewAI
Cursor
LangChain
LangGraph
LlamaIndex
LLM
RAG
AI Agents
Function Calling
Human-in-the-Loop
LLM Guardrails
OpenAI Codex
Structured Outputs
DevOps
AWS
Azure
Docker
Kubernetes
Apply
In office • Bachelor's Degree
DevOps
AWS
Azure
CI/CD
Kubernetes
Cybersecurity
FedRAMP
Apply
In office
DevOps
SLI/SLO/SLA
Management
Slack
Apply
In office
DevOps
Incident Management
Apply
In office • Bachelor's Degree
Apply
$70k – $206k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
AI Agents
Apply
$94k – $294k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Portland • Milwaukee • Dallas • Columbus
Apply
$70k – $206k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
AI Agents
Apply
$143k – $258k per year (Estimated) • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
JavaScript
Python
TypeScript
Python
pySpark
AI/ML
Prompt Engineering
Spark
DevOps
AWS
Azure
CI/CD
GCP
Git
Jenkins
GitHub
GitLab
Analytics
ETL/ELT
Apply
$163k – $434k per year • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
JavaScript
Python
TypeScript
Python
pySpark
AI/ML
Prompt Engineering
Spark
DevOps
AWS
Azure
CI/CD
GCP
Git
Jenkins
GitHub
GitLab
Analytics
ETL/ELT
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.