374,478open jobs
9,705companies
47,978added this week
Browse all
Location
In office
Seniority
Principal
Overview
Company
Impact
Profile match
Oracle is an American enterprise technology company founded in 1977 by Larry Ellison, Bob Miner and Ed Oates, and headquartered in Austin, Texas. It built its business on the Oracle Database, still the reference relational engine for large transactional systems, and has expanded into a full applications suite covering finance, human resources, supply chain and customer experience through Fusion Cloud and NetSuite. Its fastest growing segment is Oracle Cloud Infrastructure, which the company has positioned aggressively for AI training and inference workloads through large multi-year capacity contracts.

Provides technical leadership for a set of services within OCI’s network monitoring and observability platform, which provides visibility into the health, performance, and behavior of OCI network infrastructure.

Designs and evolves highly scalable, reliable, and efficient distributed services for collecting, processing, storing, and analyzing network telemetry. Owns service architecture and technical decisions across telemetry ingestion, stream processing, metrics, alerting, andnetwork health monitoring.

Ensures owned services operate reliably at cloud scale, meeting requirements for availability, data integrity, latency, scalability, and operational efficiency. Identifies performance and reliability bottlenecks and drives architectural and engineering improvements.

Serves as a technical expert for owned services, leads resolution of complex production issues, and works closely with network engineering, infrastructure, and dependent service teams. Mentors engineers and contributes to architecture and engineering standards across the broader Network Monitoring organization.

Responsibilities

Network Observability & Service Architecture

  • Own the architecture and technical evolution of critical services within OCI’s network observability platform.
  • Design distributed services that collect, process, aggregate, store, query, and expose network telemetry at cloud scale.
  • Design scalable telemetry ingestion and processing pipelines for high-volume and high-cardinality data.
  • Develop capabilities for monitoring network health, topology, state, performance, and device behavior.
  • Design solutions that enable timely detection, diagnosis, and isolation of network failures and performance degradation.
  • Ensure telemetry quality, completeness, freshness, accuracy, and availability within owned services.

Scalability & Distributed Systems

  • Design horizontally scalable and elastic services capable of supporting continued OCI infrastructure and traffic growth.
  • Identify and resolve performance, throughput, latency, storage, and scalability bottlenecks.
  • Design high-throughput streaming and event-processing systems, including partitioning, buffering, aggregation, backpressure, and failure handling.
  • Implement resilient state management, replication, synchronization, and recovery mechanisms.
  • Make appropriate engineering trade-offs across consistency, availability, latency, durability, performance, and cost.

Reliability & Operational Excellence

  • Define and meet SLOs for availability, durability, latency, data freshness, and correctness of owned services.
  • Design fault-tolerant services that operate through infrastructure failures, network disruptions, dependency failures, and software upgrades.
  • Define KPIs, telemetry, dashboards, and alerts required to understand service health and identify operational risks.
  • Lead diagnosis and resolution of complex production issues involving owned services and their dependencies.
  • Lead root-cause investigations and implement corrective actions that prevent recurrence.
  • Maintain high standards for operational readiness, capacity planning, deployment, upgrades, rollback, and recovery.
  • Participate in operational support and serve as an escalation point for critical issues involving owned services.

Network Monitoring & Analytics

  • Build capabilities for monitoring large-scale Layer 2 and Layer 3 network infrastructure.
  • Develop mechanisms to identify changes in network state, topology, reachability, performance, and device health.
  • Correlate telemetry across network devices and infrastructure services to improve fault detection and localization.
  • Develop approaches for identifying abnormal network behavior, telemetry gaps, capacity risks, and infrastructure failures.
  • Partner with network engineering teams to translate network behavior and operational requirements into monitoring capabilities.
  • Improve alert quality and signal-to-noise ratio to reduce unnecessary operational load.

Software Engineering & Automation

  • Design and implement high-quality, maintainable, and performance-sensitive production software.
  • Drive engineering practices for testing, code quality, automation, and safe software delivery within owned services.
  • Automate infrastructure provisioning, configuration, deployment, patching, upgrades, and rollback.
  • Improve service efficiency through performance optimization, capacity management, and reduction of operational toil.
  • Ensure security, compliance, and vulnerability remediation requirements are incorporated into service design and operation.

Technical Leadership & Collaboration

  • Provide technical leadership for projects and initiatives involving owned services.
  • Make and document architectural decisions and influence related designs across dependent services.
  • Lead complex technical problems that span service boundaries and coordinate with other teams when required.
  • Mentor engineers in distributed systems, networking, observability, reliability, and software engineering.
  • Conduct design and code reviews and raise engineering standards within the team.
  • Evaluate new technologies and engineering approaches where they improve scalability, reliability, performance, or operational efficiency.
  • Contribute to hiring, technical interviews, knowledge sharing, and development of engineering talent.

Skills & Technologies

The candidate should have strong expertise in several of the following areas:

  • Distributed Systems & Cloud: Distributed systems, cloud infrastructure, Kubernetes, microservices, distributed state management, high availability, fault tolerance, capacity planning, and performance engineering.
  • Observability: Prometheus, Grafana, metrics and telemetry systems, time-series data, high-cardinality metrics, alerting, dashboards, SLOs, and large-scale monitoring systems.
  • Streaming & Data Processing: Kafka, Flink or equivalent technologies; stream processing, event-driven systems, partitioning, aggregation, buffering, backpressure, and data pipelines.
  • Networking: Strong understanding of L2/L3 networking, routing and switching, BGP, LLDP, SNMP, gNMI, network topology, network failure modes, and troubleshooting.
  • Network Telemetry: Experience collecting and processing telemetry from large network environments using streaming telemetry, counters, events, protocol state, and device health information.
  • Programming: Strong software engineering skills in Java, Go, or similar languages, including concurrent and performance-sensitive production systems.
  • Reliability Engineering: SLOs, availability, durability, failure handling, load shedding, throttling, rate limiting, incident response, RCA, and production readiness.
  • Automation: Infrastructure as Code, CI/CD, automated deployment and configuration, safe rollout and rollback, and operational automation.
  • Security: Secure cloud service design, authentication and authorization, encryption, vulnerability remediation, and protection of infrastructure telemetry.

Experience building and operating large-scale cloud infrastructure, network monitoring, telemetry, or observability services is strongly preferred.

Career Level - IC4

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
374,478 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$168k – $205k per year • In office • 5+ years exp • Bachelor's Degree • New York
Java
JavaScript
Frontend
React.js
DevOps
AWS
Azure
CI/CD
GCP
Apply
$99k – $225k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Arlington
Python
Databases
Apache Kafka
ElasticSearch
RabbitMQ
Redpanda
AI/ML
NLP
DevOps
AWS
Azure
CI/CD
Docker
Kubernetes
Splunk
Cybersecurity
MISP
MITRE ATT&CK
Analytics
ETL/ELT
Apply
$46k – $99k per year (Estimated) • In office • 13+ years exp • Bengaluru
Java
Python
DevOps
AWS
Azure
GCP
IAM
Kubernetes
Platform Engineering
Rest API
Apply
$156k – $234k per year • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Irving
Java
SQL
C#
Java
Spring Boot
C#
.NET
Apply
Remote • Full-Time • 8+ years exp • Kharkiv
Bash
C++
Go
Java
Python
C
C++
OpenGL
C
FFmpeg
DevOps
Git
WebRTC
Robotics
GStreamer
Apply
$91k – $187k per year • Equity • In office • Nashville
Apply
$102k – $210k per year • Equity • In office
Apply
$93k – $210k per year • Equity • In office
Apply
$105k – $235k per year • Equity • In office • Bachelor's Degree • Redwood City
C++
Go
Java
Python
AI/ML
AI Agents
DevOps
CI/CD
Docker
Kubernetes
Rest API
Cybersecurity
Zero Trust
Robotics
Digital Twin
Gazebo
Isaac Sim
ROS
Teleoperation
Webots
IoT
MQTT
OPC UA
Apply
$105k – $235k per year • Equity • In office • Santa Clara
Apply
See all jobs
This is one of many
374,478 more open roles from verified company boards, updated every day.