368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$152k – $324k per year (Estimated)
Location
Remote/Hybrid (San Francisco, United States)
Seniority
Architect
Employment
Full-Time
Overview
Company
Impact
Profile match
Hyperbolic runs an open access cloud that aggregates idle GPUs into affordable capacity for artificial intelligence training and inference. Founded in 2023 by researchers from Berkeley and the University of Washington, it offers both raw compute rental and hosted open model endpoints. The company aims to keep frontier-scale experimentation available outside the large hyperscalers.

Who We Are

Hyperbolic Labs is on a mission to democratize AI by breaking down the barriers to computing power with our Open-Access AI Cloud. By making better use of idle computing resources across the globe, we offer an innovative GPU marketplace and AI inference service that promise affordability and accessibility for all. As pioneers at the intersection of AI and open-source technology, we believe in an open future where AI innovation is limited only by imagination, not by access to resources. We're looking for forward-thinking individuals who share our passion for making AI universally accessible, secure, and affordable. Join us in building a platform that empowers innovators everywhere to turn their visionary AI projects into reality.

About the Role

We are seeking a highly technical Vice President of Infrastructure to build and scale the foundational infrastructure powering our AI cloud platform.

This is a hands-on executive leadership role. While you will own infrastructure strategy, organizational growth, and executive-level decision making, we expect you to remain deeply engaged in architecture, design, and engineering execution. You should expect to spend approximately 30-40% of your time directly contributing to technical design, architecture reviews, debugging critical production issues, and partnering with engineers on implementation.

The ideal candidate has previously built and scaled cloud platforms, preferably GPU-native cloud infrastructure supporting AI training and inference workloads. You have experience operating at the intersection of executive leadership and hands-on engineering and are excited to help build both the technology and the team.

What You'll Own

Cloud Infrastructure Architecture

  • Lead the design and evolution of our AI cloud platform

  • Define the architecture for GPU orchestration, compute scheduling, networking, storage, and distributed systems

  • Make critical decisions regarding cloud infrastructure, bare-metal deployments, and platform scalability

  • Personally participate in architecture reviews and key technical initiatives

GPU Cloud Platform

  • Build and scale large GPU clusters supporting customer workloads

  • Design systems for GPU provisioning, scheduling, utilization optimization, and capacity management

  • Drive platform reliability and performance for AI training and inference workloads

  • Partner closely with engineering teams on infrastructure requirements for next-generation AI systems

Technical Leadership

  • Remain deeply involved in engineering decisions and technical direction

  • Contribute directly to infrastructure design and implementation efforts

  • Review architecture proposals, system designs, and major infrastructure changes

  • Act as the technical escalation point for complex infrastructure challenges

Infrastructure & Reliability

  • Establish best practices for Kubernetes, observability, CI/CD, security, and operational excellence

  • Build SRE and Platform Engineering functions from the ground up

  • Define reliability standards including SLOs, SLIs, incident response processes, and capacity planning

  • Drive automation across infrastructure operations

Organizational Leadership

  • Recruit and develop world-class Infrastructure, Platform, and SRE teams

  • Build a high-performance engineering culture focused on ownership and execution

  • Partner with executive leadership on company strategy and infrastructure investments

  • Manage infrastructure budgets, vendor relationships, and capacity planning

Required Experience

Must-Have Background

  • 12+ years building and operating large-scale infrastructure systems

  • Experience leading infrastructure organizations while remaining hands-on technically

  • Previous experience building or operating a cloud platform at scale

  • Experience building GPU infrastructure or AI/ML compute platforms

  • Proven track record scaling infrastructure in high-growth startup environments

Deep Technical Expertise

  • Expert-level Kubernetes knowledge

  • Experience designing and operating multi-region cloud infrastructure

  • Strong understanding of Linux, networking, distributed systems, and storage architecture

  • Experience with Infrastructure-as-Code and automation frameworks

  • Deep expertise in observability, monitoring, and reliability engineering

  • Experience building highly available production systems

Strongly Preferred

  • Experience with GPU scheduling, Slurm, Kubernetes GPU operators, Ray, or distributed training systems

  • Experience managing thousands of GPUs in production environments

  • Background supporting AI training and inference platforms

Hyperbolic is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$25k – $60k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Chennai
Java
PowerShell
Python
TypeScript
JavaScript
Java
Spring Boot
Databases
Amazon Redshift
DynamoDB
Frontend
Angular
DevOps
Amazon EC2
Amazon EKS
Ansible
ArgoCD
AWS
AWS Lambda
Azure
Azure DevOps
Bicep
Chef
CI/CD
CloudFormation
Configuration Management
Docker
GCP
GitLab CI
Jenkins
Kubernetes
Platform Engineering
Puppet
Terraform
Amazon S3
GitLab
IAM
Apply
$45k – $114k per year (Estimated) • Remote • Full-Time • São Paulo • Vitoria-Gasteiz • Fortaleza • Rio de Janeiro • Recife
DevOps
AWS
Azure
Azure DevOps
Bicep
CI/CD
CloudFormation
GCP
GitHub
GitHub Actions
GitLab
GitLab CI
IAM
Incident Management
Jenkins
Kubernetes
Terraform
Apply
$28k – $71k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
DevOps
CI/CD
Platform Engineering
Apply
$173k – $260k per year • In office • Full-Time • PhD • San Francisco
JavaScript
Node JS
Python
Python
Celery
Django
Flask
Databases
RabbitMQ
Redis
AI/ML
Agentforce
AI Agents
DevOps
Akamai
AWS
CI/CD
Cloudflare
CloudFormation
Helm
Jenkins
Kubernetes
Spinnaker
Terraform
Marketing
Salesforce
Apply
$45k – $127k per year (Estimated) • Remote/Hybrid • Full-Time • Recife
Java
Python
Java
Hibernate
Spring Boot
Spring Cloud
DevOps
Azure
Azure AKS
Kubernetes
Apply
In office • Full-Time
Python
AI/ML
CUDA
CUDA Toolkit
DevOps
Docker
Kubernetes
PagerDuty
SLI/SLO/SLA
SLURM
HPC
Management
Linear
Marketing
Zendesk
Apply
In office • Full-Time
AI/ML
InfiniBand
NCCL
DevOps
Ansible
Grafana
Kubernetes
Prometheus
SLURM
Terraform
Apply
$197k – $340k per year (Estimated) • In office • Full-Time • San Francisco
AI/ML
CUDA
CUDA Toolkit
InfiniBand
DevOps
CI/CD
Configuration Management
Pulumi
Terraform
Apply
$105k – $228k per year (Estimated) • In office • Full-Time • 4+ years exp • San Francisco
Python
SQL
Apply
$173k – $314k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • San Francisco
Apex
JavaScript
Node JS
Python
SQL
TypeScript
Apex
Lightning Web Components
AI/ML
Agentforce
AI Agents
Claude
Claude Code
Copilot
Cursor
LLM
RAG
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
Grafana
gRPC
Kubernetes
New Relic
Prometheus
Splunk
Marketing
Salesforce
QA
Cypress
JMeter
k6
Locust
Playwright
Postman
Rest-Assured
Selenium
Apply
Senior ML Engineer 1 hour ago
$149k – $224k per year • In office • Full-Time • 5+ years exp • Master's Degree • San Francisco • Washington • Palo Alto
Python
Python
pySpark
Databases
Apache Kafka
AI/ML
AI Agents
Agentforce
Airflow
Anomaly Detection
Feature Store
Flink
Ray
Red Teaming
Spark
DevOps
CI/CD
Docker
Kubernetes
Cybersecurity
MITRE ATT&CK
Marketing
Salesforce
Apply
In office • Internship • 1+ year exp • Bachelor's Degree • San Francisco
Go
JavaScript
Ruby
Scala
Apply
$360k – $530k per year • In office • Full-Time • Bachelor's Degree • San Francisco
MATLAB
Python
MATLAB
Simulink
AI/ML
OpenAI
Robotics
Digital Twin
Apply
$222k – $277k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • San Francisco
DevOps
CI/CD
Immutable Infrastructure
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.