368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$79k – $210k per year
Location
In office (Seattle)
Seniority
Senior
Overview
Company
Impact
Profile match
Oracle Corporation is an American multinational computer technology corporation headquartered in Austin, Texas. Founded in 1977 by Larry Ellison, Bob Miner, and Ed Oates, Oracle is one of the world's largest enterprise software and cloud computing infrastructure providers.

Designs, implements, and optimizes components in distributed systems with an emphasis on scalability, resiliency, and operability. Delivers features and load/performance tests; leverages data plane platforms and distributed state tools for high-volume retrieval, storage, and processing; and reviews peers’ implementations for scalability compliance. Builds fault-tolerant paths (redundancy, replication, automatic failover), applies recovery-oriented principles, and implements retries, circuit breakers, and timeouts. Proactively detects and mitigates issues via tests, alarms, dashboards, and telemetry; authors runbooks and participates in incident response and RCAs. Implements standard replication and synchronization, develops automation/IaC for troubleshooting and maintenance, and applies advanced security controls (encryption, access, remediation) while ensuring change, compliance, and documentation standards are met.

Oracle Cloud Infrastructure Workflow is a Tier 0 service that is critical to the smooth functioning of ALL basic OCI services by enabling their execution of distributed, multi-step work in a fault tolerant manner. An engineer on this team is responsible for the development of features making the platform more resilient and efficient including the launch of a brand new V2 version, as well as the operational excellence of the service, as we scale to meet OCI's exponential growth needs.

Key Responsibilities

System Design & Architecture - System Scalability:

-Implements and contributes to the development for components of distributed systems that support horizontal and vertical scaling including leveraging distributed state management tools.

-Optimizes code and/or systems for large-scale data processing in large-scale systems.

-Implements scalability requirements for assigned components and reviews implementation of team members.

-Leverages components of data plane platforms to handle large-scale data retrieval, storage, and processing.

-Implements performance and load testing.

System Design & Architecture - System Reliability Design:

-Collaborates with team to build fault-tolerant components capable of withstanding in-service updates by implementing redundancy, replication, and automatic failover mechanisms.

-Applies recovery oriented computing principles to design components that effectively handle service disruptions.

-Implements retry mechanisms, circuit breakers, and timeouts to help handle network unreliability.

System Design & Architecture - System Reliability Performance:

-Implements tests and alarm configurations to proactively detect and address issues/failures.

-Supports efforts to recover from failures by drafting and executing runbooks and operational procedures.

-Builds and customizes dashboards, telemetry systems, and alerting mechanisms to monitor component health.

System Design & Architecture - Correctness / Availability:

-Designs and implements functional requirements and testing for assigned features within an existing system.

-Implements tests scenarios (e.g., fault-injection, brown-out) to evaluate system correctness.

-Implements standard data replication and synchronization techniques to maintain data integrity and availability.

Operational Troubleshooting & Incident Management:

-Diagnoses, debugs, and resolves issues in system components to support ongoing operation.

-Implements basic strategies to prevent interruptions, ensuring no maintenance windows are required for customers and users when resolving issues.

-Designs and implements automation scripts and tooling used to troubleshoot operational issues.

-Participates in operational support rotations, assisting in incident responses and root cause investigations.

Compliance & Security:

-Applies advanced security measures to protect data and applications in multi-tenant environments, including encryption and access controls.

-Implements remediation plans to continuously improve security.

-Collaborates with the team to ensure cloud infrastructure complies with relevant industry standards and regulations and that documentation is up-to-date

Automation & Change Management:

-Maintains automation scripts and tools (e.g., Infrastructure as Code (IaC)) for managing cloud infrastructure.

-Adheres to change management plans for patching, updating, and rolling back applications.

Core Responsibilities

Planning & Execution:

-Track timelines with minimal supervision, ensuring work is completed in a timely manner and is in alignment with project requirements.

-Prioritize and adjust work as resources or timelines change, with some guidance

Collaboration & Partnership:

-Collaborates across teams to align on expectations and achieve shared objectives. Builds and maintains a comprehensive understanding of business, stakeholder, and/or customer needs to build and support effective partnerships. Actively listens to diverse perspectives and asks questions to ensure understanding of others.

Problem Solving:

-Independently identifies and addresses standard and non-standard issues in accordance with standard practices, escalating more complex issues as appropriate. Analyzes data and/or information from multiple sources to troubleshoot standard and non-standard errors. Contributes to knowledge sharing and best practices

Continuous Learning:

-Embraces continuous learning by actively seeking to build knowledge and new skills and/or tools, and staying current with industry trends and best practices. Seeks out and leverages feedback and training to improve skills. Contributes to a culture of continuous learning and knowledge sharing with team members..

Continuous Improvement:

-Develops ideas and recommends updates to increase the efficiency and effectiveness of processes, protocols, and workflows within a team. Seeks input from team members on alternative approaches and methods for improving work.

Disclaimer:

Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.

Range and benefit information provided in this posting are specific to the stated locations only

US: Hiring Range in USD from: $79,200 to $209,500 per annum. May be eligible for bonus and equity.

Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.

Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following:

1. Medical, dental, and vision insurance, including expert medical opinion

2. Short term disability and long term disability

3. Life insurance and AD&D

4. Supplemental life insurance (Employee/Spouse/Child)

5. Health care and dependent care Flexible Spending Accounts

6. Pre-tax commuter and parking benefits

7. 401(k) Savings and Investment Plan with company match

8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.

9. 11 paid holidays

10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.

11. Paid parental leave

12. Adoption assistance

13. Employee Stock Purchase Plan

14. Financial planning and group legal

15. Voluntary benefits including auto, homeowner and pet insurance

The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.

Career Level - IC3

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Seattle
$24k – $51k per year (Estimated) • Remote/Hybrid • Full-Time • Moscow
Node JS
JavaScript
Node JS
Commander.js
Databases
Apache Kafka
OpenSearch
PostgreSQL
Redis
DevOps
AWS
Chaos Engineering
CI/CD
GitOps
Grafana
Helm
Incident Management
Jaeger
Kubernetes
Opsgenie
PagerDuty
Prometheus
SLI/SLO/SLA
Terraform
Zabbix
Cybersecurity
Tcpdump
Management
Jira
Apply
$117k – $251k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Singapore
SQL
DevOps
Incident Management
Apply
$133k – $284k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Singapore
DevOps
Incident Management
SLI/SLO/SLA
Splunk
Management
Confluence
ServiceNow
Apply
$66k – $156k per year (Estimated) • In office • Full-Time • Singapore
SQL
DevOps
Incident Management
Apply
$66k – $162k per year (Estimated) • In office • Full-Time • 5+ years exp • Singapore
Java
COBOL
COBOL
IBM MQ
DevOps
Incident Management
SLI/SLO/SLA
Splunk
Apply
$115k – $235k per year • Equity • In office • Bachelor's Degree • Nashville
Go
Java
Python
AI/ML
AutoGen
Claude
Claude Code
Copilot
CrewAI
Cursor
LangChain
LangGraph
LlamaIndex
LLM
RAG
AI Agents
Function Calling
Human-in-the-Loop
LLM Guardrails
OpenAI Codex
Structured Outputs
DevOps
AWS
Azure
Docker
Kubernetes
Apply
In office • Bachelor's Degree
DevOps
AWS
Azure
CI/CD
Kubernetes
Cybersecurity
FedRAMP
Apply
In office
DevOps
SLI/SLO/SLA
Management
Slack
Apply
In office
DevOps
Incident Management
Apply
In office • Bachelor's Degree
Apply
$70k – $196k per year • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
Databases
Databricks
Google BigQuery
SAP HANA
Snowflake
AI/ML
Knowledge Graph
DevOps
Azure
Apply
$70k – $196k per year • Remote/Hybrid • Full-Time • 5+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
DevOps
SLI/SLO/SLA
Apply
$70k – $206k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
AI Agents
Apply
$88k – $146k per year • In office • Full-Time • 3+ years exp • Master's Degree • Seattle
Perl
Python
Ruby
SQL
Databases
MySQL
Teradata
DevOps
AWS
Analytics
Tableau
Apply
$180k – $225k per year • Equity • In office • Full-Time • 4+ years exp • Bachelor's Degree • Seattle
C#
Go
Java
Kotlin
TypeScript
JavaScript
AI/ML
Claude
Claude Code
OpenAI Codex
Frontend
React.js
DevOps
AWS
AWS Step Functions
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.