678,285open jobs
39,280companies
100,060added this week
Browse all
Salary
$255k – $490k per year
Location
In office (San Francisco)
Employment
Full-Time
Overview
Company
Impact
Profile match
OpenAI is an American artificial intelligence research and deployment company founded in 2015 with the mission of ensuring that artificial general intelligence benefits all of humanity. It develops the GPT family of large language models and turns them into consumer and developer products, including the ChatGPT assistant, the Sora video model, the Codex coding agent and a commercial API used by millions of developers. Structured as a public benefit corporation controlled by a non-profit foundation, the company is backed by Microsoft and SoftBank and operates from San Francisco.

About the Team

Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. Our systems turn large, heterogeneous fleets of machines into dependable compute for research and products.

We build Kubernetes-based control planes, controllers, services, and APIs that coordinate the lifecycle of machines and clusters. We connect global infrastructure management with the realities of bare-metal systems, giving clients consistent interfaces across differences in hardware, topology, and provider behavior.

About the Role

You will build distributed systems that provision, configure, and manage compute throughout its lifecycle. Your work will connect global services and Kubernetes controllers with the systems that bring machines online, update them safely, and recover them when something goes wrong.

This role combines software architecture with an understanding of how machines and data centers work. You might design a lifecycle API, improve controller performance under high concurrency and provider rate limits, or trace a provisioning failure from an API through reconciliation to network boot or host configuration. You will help these systems remain reliable as the fleet expands across sites and generations of GPU hardware.

We value depth in relevant systems and the ability to connect layers. You do not need to arrive as an expert in every component of the stack.

In this role, you will:

  • Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows.

  • Define APIs and resource models that let clients request and track lifecycle operations through consistent interfaces across hardware platforms and providers.

  • Build provisioning and configuration services that coordinate network boot, hardware management interfaces, and the deployment of firmware, operating-system images, drivers, and host configuration.

  • Develop lifecycle management for discovery, allocation, provisioning, upgrades, maintenance, recovery, and decommissioning, integrating with health and validation systems.

  • Design reliable reconciliation and recovery through concurrent changes, interrupted operations, and partial failures, with staged rollouts that limit disruption across nodes, racks, and clusters.

  • Improve control-plane throughput, API latency, and the time infrastructure takes to reach its desired state, while respecting the limits of site systems and provider APIs.

  • Build the software integrations that bring new sites and GPU hardware generations into the platform, partnering with hardware, networking, data-center, and other infrastructure teams.

You might thrive in this role if you:

  • Have strong software engineering fundamentals and experience designing, implementing, and owning production distributed systems or infrastructure services.

  • Have experience developing infrastructure systems that use Kubernetes APIs and reconciliation to manage resources.

  • Understand how a bare-metal node moves from power-on to a configured, workload-ready system, with depth in one or more areas such as PXE, DHCP/DNS, baseboard management controllers (BMCs), firmware, Linux, drivers, images, or configuration management.

  • Can design reliable APIs and asynchronous workflows, reasoning about concurrency, consistency, idempotency, and failures across service and provider boundaries.

  • Can diagnose reliability and performance problems across service, operating-system, and machine boundaries, turning production evidence into lasting software improvements.

  • Work effectively across engineering specialties and communicate system behavior and technical tradeoffs clearly.

Bonus points if you:

  • Have built infrastructure control planes that coordinate operations across multiple sites or regions.

  • Have worked with GPU or HPC infrastructure, including topology and shared dependencies across machines, racks, or clusters.

  • Have integrated multiple hardware platforms or infrastructure providers into a common service or resource model.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.

We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement.

Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.

To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link.

OpenAI Global Applicant Privacy Policy

At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
678,285 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$130k per year • Remote/Hybrid • TS/SCI • 2+ years exp • Bachelor's Degree
Python
JavaScript
Java
TypeScript
Java
Spring Boot
Databases
ElasticSearch
Frontend
Vue.js
Angular
React.js
DevOps
Terraform
Ansible
GitLab CI
CI/CD
Git
AWS
Docker
Kubernetes
Analytics
Apache NiFi
Apply
Counsel, Commercial 3 hours ago
$252k – $280k per year • Remote/Hybrid • Full-Time • 7+ years exp • San Francisco
AI/ML
OpenAI
Apply
$35k – $76k per year (Estimated) • Remote • Full-Time • Amsterdam
DevOps
Red Hat
Kubernetes
Apply
$18k – $50k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Centurion
Python
SQL
PowerShell
Databases
MySQL
DevOps
Podman
CI/CD
Jenkins
Docker
Kubernetes
Bitbucket
Management
Jira
Agile
Scrum
Apply
$55k – $80k per year • In office • Full-Time • 6+ years exp • Sandton
JavaScript
TypeScript
C#
C#
ASP.NET Core
Entity Framework Core
Blazor
WPF
xUnit
Moq
Databases
MS SQL
Azure Cosmos DB
Frontend
Angular
DevOps
Azure DevOps
Azure
Git
Docker
Kubernetes
Management
Agile
Scrum
Apply
Counsel, Commercial 3 hours ago
$252k – $280k per year • Remote/Hybrid • Full-Time • 7+ years exp • San Francisco
AI/ML
OpenAI
Apply
$272k – $302k per year • Remote/Hybrid • Full-Time • 5+ years exp • San Francisco
AI/ML
OpenAI
Apply
$266k – $455k per year • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • San Francisco
AI/ML
OpenAI
Apply
$128k – $282k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • Munich
AI/ML
OpenAI
Apply
$110k – $295k per year (Estimated) • Remote/Hybrid • Full-Time • Dublin
AI/ML
OpenAI
DevOps
SLI/SLO/SLA
Apply
$73k – $133k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree • Los Angeles • San Jose • San Francisco • San Diego • Bellevue
Design
AutoCAD
Apply
$148k – $222k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • San Francisco
Apply
$70k – $111k per year • In office • Full-Time • 6+ years exp • Bachelor's Degree • Minneapolis • Seattle • San Francisco • Portland
Cybersecurity
Wireshark
Apply
Technical Recruiter 8 hours ago
$200k – $275k per year • In office • Full-Time • San Francisco
DevOps
Kubernetes
Apply
$79k – $188k per year (Estimated) • Remote • Full-Time • 3+ years exp • San Francisco
JavaScript
Frontend
React.js
Apply
See all jobs
This is one of many
678,285 more open roles from verified company boards, updated every day.