600,580open jobs
28,220companies
85,699added this week
Browse all
Location
In office
Seniority
Senior · 4+ years exp
Overview
Company
Impact
Profile match
Oracle is an American enterprise technology company founded in 1977 by Larry Ellison, Bob Miner and Ed Oates, and headquartered in Austin, Texas. It built its business on the Oracle Database, still the reference relational engine for large transactional systems, and has expanded into a full applications suite covering finance, human resources, supply chain and customer experience through Fusion Cloud and NetSuite. Its fastest growing segment is Oracle Cloud Infrastructure, which the company has positioned aggressively for AI training and inference workloads through large multi-year capacity contracts.

Job Responsibilities

  • Improve the reliability, scalability, performance, and operational efficiency of assigned OCI Compute services and components.
  • Investigate and resolve complex production incidents; contribute to mitigation, recovery, RCA, and follow-up actions.
  • Own and improve service-level KPIs, SLOs, dashboards, alerting, deployment validation, and operational procedures for assigned systems.
  • Build automation and tooling to reduce operational toil and improve production safety.
  • Partner with development and infrastructure teams on service architecture, deployment, configuration, and reliability improvements.
  • Use observability, telemetry, event correlation, and AIOps capabilities to improve detection, diagnosis, and incident response.
  • Support upgrades, migrations, patching, capacity planning, performance tuning, security vulnerability managementand production rollouts.
  • Troubleshoot distributed-system issues by analyzing service topology, dependencies, configuration, and failure modes.
  • Contribute to incident-management practices, operational readiness, and service ownership improvements.
  • Share technical knowledge and support team members through documentation, reviews, and collaboration.
  • Participate in a 12x7 on-call rotation and support response to customer-impacting incidents.

Mandatory Skills

  • 4-8 years of experience in SRE, Production Engineering, Cloud Operations, Systems Engineering, or a similar role.
  • Experience operating and improving highly available production systems.
  • Strong programming or scripting skills in Python, Java, Go, or similar languages.
  • Hands-on experience with Linux, cloud infrastructure, networking, compute, and storage.
  • Experience with production monitoring, alerting, dashboards, logs, metrics, and tracing.
  • Experience owning or improving service SLIs, SLOs, KPIs, and operational procedures.
  • Strong incident troubleshooting, RCA, debugging, and problem-solving skills.
  • Experience with deployment pipelines, release validation, automation, and change-management practices.
  • Understanding of distributed systems, service dependencies, capacity planning, and performance tuning.
  • Ability to work independently on technical problems and collaborate effectively with engineering teams.
  • Strong written and verbal communication skills.

Preferred Skills

  • Experience with OCI and cloud infrastructure services.
  • Experience with AIOps, anomaly detection, event correlation, predictive alerting, or automated remediation.
  • Experience with Kubernetes, containers, infrastructure-as-code, and CI/CD.
  • Experience with service migrations, fleet maintenance, upgrades, patching, or production rollouts.
  • Experience with architecture reviews, operational-readiness reviews, and post-incident improvements.
  • Experience contributing to technical initiatives, knowledge sharing, code reviews, or operational improvements within the team.
  • Familiarity with security, compliance, and access-control practices in production environments.

Self-Test Questions

  • Do you have 4-8 years of relevant SRE, Production Engineering, Cloud Operations, or Systems Engineering experience?
  • Have you independently operated or improved a production service, system, or infrastructure component?
  • Can you investigate production incidents and contribute to mitigation, recovery, RCA, and follow-up actions?
  • Do you have hands-on experience with Linux, cloud infrastructure, distributed systems, networking, compute, or storage?
  • Are you proficient in Python, Java, Go, or a similar language for automation, tooling, and troubleshooting?
  • Have you built or improved automation, deployment validation, CI/CD pipelines, or operational tooling?
  • Do you have experience with monitoring, alerting, logs, metrics, tracing, and service health indicators such as SLOs or KPIs?
  • Can you work independently on assigned technical problems, collaborate with partner teams, and participate in a 12x7 on-call rotation?

Career Level - IC3

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
600,580 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$130k – $196k per year • Equity • In office • Full-Time • 5+ years exp • Bachelor's Degree • Boulder • Atlanta
Python
Java
TypeScript
Java
Spring Boot
AI/ML
Cursor
Claude Code
AI Agents
DevOps
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Management
Agile
Apply
$123k – $126k per year • Remote • Full-Time • Poland
Go
Java
TypeScript
SQL
Databases
Oracle
Apache Kafka
Apply
$116k – $160k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Austin
Python
Cybersecurity
CAPA
Apply
$36k – $55k per year (Estimated) • In office • Internship • Bachelor's Degree • Indianapolis
Python
Chips/EDA
Altium Designer
Apply
$72k – $108k per year • In office • Full-Time • Bachelor's Degree • United States
Python
SQL
MATLAB
Apply
$62k – $100k per year • Equity • In office • 3+ years exp • Austin
Management
Agile
Apply
Remote/Hybrid • PhD
Java
Java
GraalVM
Databases
Oracle
Apply
In office
DevOps
Incident Management
Apply
$147k – $244k per year • Equity • In office • 5+ years exp
Apply
Remote/Hybrid
Apply
See all jobs
This is one of many
600,580 more open roles from verified company boards, updated every day.