433,514open jobs
15,105companies
61,322added this week
Browse all
Salary
$138k – $345k per year (Estimated)
Location
In office (San Francisco)
Employment
Full-Time
Overview
Company
Impact
Profile match
Hyperbolic runs an open-access AI cloud that aggregates idle GPU capacity from data centres and independent providers into affordable compute for model training and inference. Developers rent hardware by the hour or call hosted open-source models through an API, at prices well below the large clouds because that supply would otherwise sit unused between other customers' workloads. The company was founded by machine learning researchers and publishes tooling for verifying that rented hardware really performed the work it charged for.

Who We Are

Hyperbolic Labs is on a mission to democratize AI by breaking down the barriers to computing power with our Open-Access AI Cloud. By making better use of idle computing resources across the globe, we offer an innovative GPU marketplace and AI inference service that promise affordability and accessibility for all. As pioneers at the intersection of AI and open-source technology, we believe in an open future where AI innovation is limited only by imagination, not by access to resources. We're looking for forward-thinking individuals who share our passion for making AI universally accessible, secure, and affordable. Join us in building a platform that empowers innovators everywhere to turn their visionary AI projects into reality.

About the Role

We're looking for an Inference Engineer to build inference capabilities on top of Forge, our unified control plane, so customers can consume model tokens without managing GPUs and our NeoCloud partners get a full-stack path to their own token-factory offering. You'll own how models get deployed and served across clusters distributed around the world, on heterogeneous hardware.

Deployment comes first: serving models on Forge and our Kubernetes offering, evaluating inference frameworks, and standing up the monitoring, gateways, and endpoints that make a deployment production-ready. From there the work expands into optimization, autoscaling, KV-cache orchestration, and customer inference debugging. This is the primary seat for inference at Hyperbolic - you'll build it end to end, with real influence over where the scope lands.

Who You Are

  • Strong general inference background with a broad, high-level command of the stack rather than a narrow specialty - you can reason about the whole path from request to token

  • Deep Kubernetes experience, including hands-on ability to operate clusters in production, not just deploy to them

  • Solid grasp of the concepts that govern inference performance: TTFT, disaggregated inference, speculative decoding, and KV cache and its inner workings

  • Familiarity with modern inference frameworks and serving engines, and the judgment to evaluate and select among them for a given workload

  • Working knowledge of NVIDIA Dynamo and how it fits into a distributed serving architecture

  • Experience setting up monitoring, gateways, and endpoints for production inference services

  • Proven ability to build a product end to end - you've taken something from nothing to serving real traffic

  • Strong self-initiative and comfort operating as the primary owner of an area with minimal direction

  • Generalist instincts: you're willing to pick up adjacent work when it's what the product needs

Preferred Qualifications

  • Experience spanning both inference deployment and inference optimization

  • Hands-on model optimization work - quantization, batching strategies, kernel-level tuning, or similar

  • Understanding of RDMA and high-performance networking as they apply to distributed serving

  • Experience deploying inference across heterogeneous accelerators

  • Background supporting customers directly on inference debugging and performance issues

  • Experience at a GPU cloud, inference provider, or AI infrastructure company

Hyperbolic is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
433,514 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$120k – $286k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Singapore
Python
JavaScript
TypeScript
SQL
Python
FastAPI
AI/ML
Quantization
Prompt Engineering
Function Calling
AI Agents
LLM
RAG
Streamlit
Anomaly Detection
Gaussian Splatting
LLMOps
Agentic Workflows
Tool Use
Frontend
Angular
React.js
DevOps
GCP
OpenShift
Azure
CI/CD
AWS
Docker
Kubernetes
Game Dev
Unity
Unreal Engine
Robotics
ROS
Gazebo
NVIDIA Omniverse
Isaac Sim
SLAM
Sim-to-Real
Imitation Learning
Digital Twin
Apply
AI Engineer, SMAI 9 hours ago
In office • Full-Time • 2+ years exp • Bachelor's Degree • Singapore
Python
JavaScript
TypeScript
SQL
Python
FastAPI
AI/ML
Prompt Engineering
AI Agents
LLM
RAG
Streamlit
Frontend
Angular
React.js
DevOps
GCP
OpenShift
GitHub Actions
Azure
CI/CD
AWS
Docker
Kubernetes
GitHub
Apply
$36k – $90k per year (Estimated) • In office • Full-Time • 5+ years exp • Pune
Python
Java
SQL
Python
Django
Java
Gradle
Databases
MySQL
PostgreSQL
Redis
Databricks
RabbitMQ
Apache Kafka
Azure SQL Database
AI/ML
Pandas
NumPy
LLM
Edge AI
DevOps
Azure DevOps
GitHub Actions
Azure
CI/CD
Jenkins
Git
Docker
Kubernetes
Azure AKS
GitHub
QA
Swagger
Apply
AI / ML Engineer 9 hours ago
$35k – $89k per year (Estimated) • In office • Full-Time • 5+ years exp • Pune
Python
SQL
Databases
Databricks
Delta Lake
Apache Kafka
AI/ML
LangGraph
LangChain
Spark
Embeddings
Great Expectations
CrewAI
RAG
OpenAI
Google AI Studio
DevOps
Azure DevOps
GitHub Actions
Azure
CI/CD
Docker
Kubernetes
Vector
GitHub
Analytics
Azure Data Factory
Apply
$89k – $176k per year (Estimated) • Remote • Full-Time
DevOps
Terraform
GCP
CloudFormation
CI/CD
AWS
Kubernetes
IAM
Cybersecurity
ISO 27001
SOC 2
Zero Trust
Apply
$106k – $261k per year (Estimated) • In office • Full-Time • San Francisco
Python
Apply
Head of Marketing 1 day ago
$136k – $289k per year (Estimated) • In office • Full-Time • 8+ years exp • San Francisco
Apply
GTM (Operations) 7 days ago
In office • Full-Time
Management
QuickBooks
Marketing
HubSpot
Apply
In office • Full-Time
Python
AI/ML
CUDA Toolkit
CUDA
DevOps
PagerDuty
SLURM
Docker
Kubernetes
SLI/SLO/SLA
HPC
Management
Linear
Apply
Capital Markets Lead 14 days ago
$112k – $249k per year (Estimated) • In office • Full-Time • 5+ years exp • San Francisco
Apply
$255k – $330k per year • In office • Full-Time • 15+ years exp • Bachelor's Degree • San Francisco
AI/ML
Supervision
Apply
$260k – $302k per year • Remote/Hybrid • Full-Time • 7+ years exp • San Francisco
Python
AI/ML
LLM
OpenAI
Apply
$401k – $445k per year • In office • Full-Time • San Francisco
AI/ML
ChatGPT
OpenAI
Post-training
LLM Guardrails
DevOps
Platform Engineering
Apply
$82k – $123k per year • In office • Full-Time • 2+ years exp • PhD • San Francisco • Chicago • New York
AI/ML
Claude
Claude Code
AI Agents
Agentforce
Apply
$82k – $123k per year • In office • Full-Time • 2+ years exp • PhD • San Francisco • Chicago • Boston • New York • Denver
AI/ML
Claude
AI Agents
OpenAI
Agentforce
Analytics
Tableau
Management
Slack
Google Workspace
Google Sheets
Apply
See all jobs
This is one of many
433,514 more open roles from verified company boards, updated every day.