368,634open jobs
9,437companies
50,578added this week
Browse all
Location
In office
Employment
Full-Time
Overview
Company
Impact
Profile match
Hyperbolic runs an open access cloud that aggregates idle GPUs into affordable capacity for artificial intelligence training and inference. Founded in 2023 by researchers from Berkeley and the University of Washington, it offers both raw compute rental and hosted open model endpoints. The company aims to keep frontier-scale experimentation available outside the large hyperscalers.

Who We Are

Hyperbolic Labs is on a mission to democratize AI by breaking down the barriers to computing power with our Open-Access AI Cloud. By making better use of idle computing resources across the globe, we offer an innovative GPU marketplace and AI inference service that promise affordability and accessibility for all. As pioneers at the intersection of AI and open-source technology, we believe in an open future where AI innovation is limited only by imagination, not by access to resources. We're looking for forward-thinking individuals who share our passion for making AI universally accessible, secure, and affordable. Join us in building a platform that empowers innovators everywhere to turn their visionary AI projects into reality.

About the Role

When a customer's GPU cluster has a problem, you are the first engineer they talk to, and most of the time you are the only one they need to. You own the ticket from the first response to the resolution, you own the clock against our SLA, and you keep ownership even when the problem needs deeper help.

This is not a ticket routing job. You will be in Linux every day, on real infrastructure, fixing real customer problems. The people who do this well here move into the Forward Deployed Engineer role, and we will build that path with you deliberately.

Who You Are

  • Ticket ownership, end to end. You own every ticket you pick up, including after it escalates. You do not hand off, you pull in the engineer you need and stay on it until the customer is working.

  • The SLA clock. First response, severity classification, and keeping us honest against our response commitments. You are the person who knows where every open issue stands.

  • First technical response and triage. Reproduce the problem, gather the logs and configuration that matter, and make the first call on whether the fault is ours or the provider's.

  • Customer environment access and configuration. SSH key and access issues, NFS mounts and storage, quotas, security groups, container and driver questions, billing and account questions.

  • Runbook execution and authoring. Run the documented play when there is one. When you solve something new, write the runbook so the next person does not escalate it.

  • Documentation. Keep our customer-facing docs and internal knowledge base current. Most repeat tickets are a documentation gap.

  • Very strong Linux experience and daily work in the CLI.

  • Experience owning tickets against a response SLA in cloud, hosting, or infrastructure support.

  • Solid networking and storage fundamentals: SSH, NFS and mounts, DNS, firewalls and security groups.

  • Working familiarity with GPU workloads: nvidia-smi, drivers, CUDA, containers.

  • Clear and fast written communication under time pressure. Customers read what you write while they are blocked.

  • Good judgment about the limits of your own knowledge, and a bias toward escalating early with a complete picture rather than late with a guess.

  • Comfortable working across time zones and with an on-call rotation for critical issues.

Preferred Qualifications

  • Experience with ticketing and on-call tooling (Zendesk, Linear, PagerDuty, or similar).

  • Scripting in bash or Python to automate repeat work.

  • Exposure to Slurm, Kubernetes, or Docker in a multi-tenant environment.

  • Background in GPU cloud, HPC, or a hardware-adjacent support org.

Hyperbolic is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$14k – $39k per year (Estimated) • In office • Full-Time • Saint Petersburg
SQL
DevOps
SLI/SLO/SLA
Cybersecurity
ISO 27001
Apply
$34k – $83k per year (Estimated) • Remote/Hybrid • Full-Time • Bengaluru
Node JS
JavaScript
Databases
Apache Kafka
DevOps
ArgoCD
Azure
Azure AKS
Azure DevOps
CI/CD
Datadog
FinOps
GitHub Actions
Grafana
Istio
Kubernetes
Platform Engineering
Prometheus
SLI/SLO/SLA
Terraform
GitHub
IAM
Cybersecurity
GDPR
Microsoft Defender
Microsoft Defender for Cloud
Okta
PCI DSS
Management
ServiceNow
Apply
$32k – $64k per year (Estimated) • Remote • Contractor • Saint Petersburg
Python
DevOps
CI/CD
GitLab CI
Kubernetes
SLI/SLO/SLA
GitLab
Apply
Remote/Hybrid • Full-Time • Riga
Go
Python
Rust
Databases
ElasticSearch
OpenSearch
AI/ML
Anomaly Detection
DevOps
AIOps
AWS
Cortex
eBPF
Grafana
Incident Management
Jaeger
Kubernetes
Loki
Mimir
OpenTelemetry
Platform Engineering
Prometheus
SLI/SLO/SLA
Splunk
Thanos
HPC
Apply
Remote/Hybrid • Full-Time • Sofia
Go
Python
Rust
Databases
ElasticSearch
OpenSearch
AI/ML
Anomaly Detection
DevOps
AIOps
AWS
Cortex
eBPF
Grafana
Incident Management
Jaeger
Kubernetes
Loki
Mimir
OpenTelemetry
Platform Engineering
Prometheus
SLI/SLO/SLA
Splunk
Thanos
HPC
Apply
VP of Engineering 2 months ago
$152k – $324k per year (Estimated) • Remote/Hybrid • Full-Time • San Francisco
AI/ML
Ray
DevOps
CI/CD
Kubernetes
Platform Engineering
SLURM
Apply
In office • Full-Time
AI/ML
InfiniBand
NCCL
DevOps
Ansible
Grafana
Kubernetes
Prometheus
SLURM
Terraform
Apply
$197k – $340k per year (Estimated) • In office • Full-Time • San Francisco
AI/ML
CUDA
CUDA Toolkit
InfiniBand
DevOps
CI/CD
Configuration Management
Pulumi
Terraform
Apply
$105k – $228k per year (Estimated) • In office • Full-Time • 4+ years exp • San Francisco
Python
SQL
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.