997,024open jobs
59,481companies
165,643added this week
Browse all
Salary
≈ $185k – $333k per year (Estimated)
Location
Hybrid (Palo Alto, United States)
Seniority
Staff · 5+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 1, 2026. First seen by Alion on Oct 27, 2025. Zettabyte scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Zettabyte builds next-generation AI data centers with high-efficiency GPU infrastructure, liquid-cooling systems, and sovereign-ready AIDC software.

About Us

At Zettabyte, we’re on a mission to make AI compute ubiquitous, seamless, and limitless. We’re building a cloud where AI just works-anywhere, anytime. “AI Power. Everywhere.” Be part of the team designing the infrastructure for the AI-first world.

Why this role exists

We need a Backend Engineer to build the systems that orchestrate GPU clusters for AI workloads. You'll create APIs that handle GPU allocation, memory management, compute scheduling, and multi-tenant isolation-challenges unique to AI infrastructure that go far beyond typical backend engineering. As part of our backend team, you'll solve problems like: How do we efficiently share expensive GPU resources across users? How do we handle GPU memory constraints for large AI models? How do we ensure quality of service when workloads compete for compute? This is an opportunity to build infrastructure where every API call could allocate thousands of dollars worth of compute per hour, where your optimizations directly impact whether AI startups can afford to train their models.

What you’ll do

  • Design APIs that abstract complex GPU operations into simple developer experiences

  • Build scheduling algorithms that maximize GPU utilization while ensuring SLA compliance

  • Develop resource management systems for GPU lifecycle-provisioning, allocation, scheduling, and release

  • Create usage tracking and billing systems for GPU-hours, memory usage, and compute utilization

  • Implement monitoring for GPU-specific metrics, health checks, and automatic failure recovery

  • Build multi-tenancy systems with resource isolation, quota management, and fair scheduling

  • Optimize cold starts for model serving and implement efficient model loading strategies

  • Collaborate with frontend engineers to expose complex infrastructure through intuitive interfaces

  • Leverage AI-assisted coding tools (GitHub Copilot, Claude Code, Cursor IDE, etc.) to boost productivity and code quality.

You’ll thrive here if you

  • 5+ years backend engineering experience with distributed systems

  • Strong proficiency in Go, Python, or similar backend languages

  • Experience with resource scheduling, orchestration, and API design (REST, GraphQL, gRPC)

  • Understanding of hardware constraints and system optimization

  • Linux systems knowledge and containerization experience (Docker, Kubernetes)

  • Comfortable working with expensive resources where efficiency directly impacts costs

  • Excited about solving novel problems in AI infrastructure (not just another CRUD app)

  • Startup mindset-comfortable with ambiguity and rapid iteration

Bonus qualifications

  • GPU or HPC cluster management experience

  • Understanding of ML/AI workload patterns and requirements

  • Experience with high-value resource allocation systems

  • Background in performance optimization for compute-intensive workloads

  • Familiarity with GPU virtualization and sharing technologies

  • Experience building billing or metering systems

Details

  • We provide Competitive salary and equity based on your experience and skillset;

  • This is a Hybrid role - 3 days in office, 2 days WFH; Must locate in Palo Alto

  • Applicants must be authorized to work in the United States without need for visa sponsorship.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
997,024 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
Palo Alto
≈ $160k – $289k per year (Estimated) • Hybrid • Full-Time • United States
JavaScript
TypeScript
AI/ML
AI Agents
Google AI Studio
Frontend
React.js
DevOps
WebSockets
CI/CD
Design
Figma
Management
Miro
Apply
Software Engineer, Sr 4 months ago
≈ $123k – $228k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Newberg
Python
TypeScript
C++
C++
Qt
DevOps
RTOS
Linux
Management
Agile
Apply
≈ $150k – $271k per year (Estimated) • Hybrid • Full-Time • 15+ years exp • Bachelor's Degree • Atlanta
Python
Go
JavaScript
TypeScript
DevOps
GCP
Azure
CI/CD
AWS
Management
Agile
Apply
≈ $130k – $257k per year (Estimated) • In office • 5+ years exp • Houston
JavaScript
Java
Kotlin
TypeScript
SQL
Java
Maven
Spring Boot
Kotlin
Mockito
Databases
PostgreSQL
Apache Kafka
AI/ML
Copilot
Claude Code
AI Agents
Agentic Workflows
Frontend
React.js
DevOps
Rest API
Splunk
CI/CD
Jenkins
Git
Grafana
AppDynamics
Cybersecurity
SonarQube
Fortify
Management
Agile
QA
Playwright
Apply
$195k per year • Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Orange
JavaScript
SQL
C#
C#
ASP.NET Core
Databases
MS SQL
Frontend
React.js
JQuery
DevOps
Rest API
GCP
Azure
CI/CD
AWS
Analytics
SSIS
Management
Scrum
Apply
In office • Master's Degree
Python
Management
Microsoft Office
Apply
≈ $7k – $18k per year (Estimated) • In office • Full-Time • 1+ year exp • Bachelor's Degree • India
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
AI/ML
Spark
DevOps
Terraform
GCP
Azure DevOps
GitHub Actions
Azure
CI/CD
Git
AWS
Platform Engineering
Incident Management
GitHub
Apply
≈ $12k – $24k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • India
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
AI/ML
Spark
DevOps
Terraform
GCP
Azure DevOps
GitHub Actions
Azure
CI/CD
Git
AWS
Platform Engineering
Incident Management
GitHub
Apply
≈ $18k – $48k per year (Estimated) • Hybrid • Full-Time • 6+ years exp • Bengaluru
Python
SQL
Databases
Amazon Redshift
AI/ML
Claude
LLM
OpenAI
DevOps
CI/CD
AWS
AWS Lambda
Amazon S3
Analytics
ETL/ELT
Management
Agile
Scrum
Apply
≈ $60k – $123k per year (Estimated) • In office • Full-Time • Master's Degree • Singapore
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
OpenCV
Quantization
Knowledge Distillation
Computer Vision
ONNX
TensorFlow
PyTorch
TPU
Edge AI
Model Distillation
Robotics
Localization
Apply
≈ $107k – $215k per year (Estimated) • Hybrid • Full-Time • Palo Alto
Go
JavaScript
TypeScript
Databases
NATS
RabbitMQ
Apache Kafka
AI/ML
Cursor
Claude Code
OpenAI Codex
Frontend
Vue.js
Next.js
React.js
DevOps
gRPC
Terraform
GCP
Datadog
Prometheus
Pulumi
Azure
CI/CD
AWS
Kubernetes
Grafana
Linux
DNS
Design
Figma
Apply
Hybrid • Internship • Bachelor's Degree • Palo Alto
Python
Go
C++
AI/ML
Copilot
Cursor
ChatGPT
DevOps
GCP
Prometheus
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Grafana
GitHub
Linux
Unix
Apply
≈ $50k – $85k per year (Estimated) • Hybrid • Internship • Bachelor's Degree • Palo Alto
Python
SQL
AI/ML
Cursor
Claude
ChatGPT
Machine Learning
Management
Google Sheets
Apply
Head of People 7 months ago
≈ $154k – $302k per year (Estimated) • Hybrid • Full-Time • Bachelor's Degree • Palo Alto
Apply
≈ $154k – $319k per year (Estimated) • Hybrid • Full-Time • 7+ years exp • Palo Alto
Python
AI/ML
InfiniBand
Red Teaming
DevOps
GCP
Cilium
Azure
CI/CD
AWS
Kubernetes
eBPF
HPC
Cybersecurity
Falco
ISO 27001
SOC 2
Zero Trust
Threat Modeling
Calico
Cryptography
Vault
Apply
$193k – $262k per year • Equity • In office • Full-Time • 5+ years exp • Bachelor's Degree • Palo Alto
AI/ML
AI Agents
DevOps
Platform Engineering
Management
Agile
Apply
≈ $163k – $324k per year (Estimated) • In office • 7+ years exp • Palo Alto
Python
Java
Databases
RabbitMQ
Apache Kafka
DevOps
Splunk
Terraform
OpenTelemetry
Datadog
Dynatrace
Prometheus
CI/CD
GitOps
Grafana
SRE
AppDynamics
AIOps
Incident Management
Error Budget
SLI/SLO/SLA
Windows
Web3
IPFS
Apply
$42k per year • In office • Palo Alto
Apply
$38k – $40k per year • In office • PhD • Palo Alto
Apply
$210k – $284k per year • Equity • In office • Full-Time • Bachelor's Degree • Palo Alto
AI/ML
AI Agents
AWS Bedrock
LLM
RAG
AWS Bedrock AgentCore
LLM Guardrails
DevOps
SLURM
AWS
HPC
Apply
See all jobs
This is one of many
997,024 more open roles from verified company boards, updated every day.