368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$200k – $550k per year
Location
In office (San Francisco)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Magic is a San Francisco research company founded in 2022 that trains frontier models for software engineering. It focuses on extremely long context windows so a model can hold entire codebases and their history in working memory while making changes. The company raised large rounds from investors including Alphabet's CapitalG and Nat Friedman and builds its own training infrastructure.

Magic’s mission is to build safe AGI that accelerates humanity’s progress on the world’s most important problems. We believe the most promising path to safe AGI lies in automating research and code generation to improve models and solve alignment more reliably than humans can alone. Our approach combines frontier-scale pre-training, domain-specific RL, ultra-long context, and inference-time compute to achieve this goal.

About the role

As an engineer on the Supercomputing Platform & Infrastructure team, you will design, build, and operate the large-scale GPU infrastructure that powers Magic’s model training and inference workloads.

A core part of this role is building and maintaining our infrastructure using Terraform-driven infrastructure-as-code practices, ensuring reproducibility, reliability, and operational clarity across clusters spanning thousands of GPUs.

Magic’s long-context models create sustained pressure on compute, networking, and storage systems. Long-running distributed jobs, high-throughput data movement, and strict availability requirements demand infrastructure that is automated, observable, and resilient by design. You will own the systems and IaC foundations that make this possible, including the Kubernetes (K8s) environments that coordinate workloads across our GPU infrastructure.

This role can evolve into broader ownership of supercomputing platform architecture, shaping how Magic scales GPU clusters and infrastructure reliability as model workloads grow.

What you’ll work on

  • Design and operate large-scale GPU clusters for training and inference

  • Build and maintain infrastructure using Terraform across cloud and hybrid environments

  • Deploy, operate, and optimize K8s clusters used to schedule and manage AI workloads

  • Develop modular, scalable IaC patterns for compute, networking, and storage provisioning

  • Improve deployment reproducibility, environment consistency, and operational safety

  • Optimize networking and storage systems for high-throughput AI workloads

  • Automate fault detection and recovery across distributed clusters

  • Debug complex cross-layer issues spanning hardware, drivers, networking, storage, OS, and cloud

  • Improve observability, monitoring, and reliability of core platform systems

What we’re looking for

  • Strong software engineering skills with experience building production infra systems

  • Deep, hands-on experience with Terraform, including module design, state management, environment isolation, and large-scale deployments

  • Experience operating production GPU infrastructure or high-performance distributed systems

  • Strong understanding of networking and storage systems

  • Experience with major cloud platforms (GCP, AWS, Azure, OCI, etc.)

  • Track record of owning production-critical infrastructure end-to-end

Our culture

  • Integrity. Words and actions should be aligned

  • Hands-on. At Magic, everyone is building

  • Teamwork. We move as one team, not N individuals

  • Focus. Safely deploy AGI. Everything else is noise

  • Quality. Magic should feel like magic

Magic strives to be the place where high-potential individuals can do their best work. We value quick learning and grit just as much as skill and experience.

Compensation, benefits, and perks (US):

  • Annual salary range between $200K - $550K depending on experience

  • Equity is a significant part of total compensation, in addition to salary

  • 401(k) plan with 6% salary matching

  • Generous health, dental and vision insurance for you and your dependents

  • Unlimited paid time off

  • Visa sponsorship and relocation stipend to bring you to SF, if possible

  • A small, fast-paced, highly focused team

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$140k – $225k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree • Seattle
C++
Go
Python
Rust
Python
FastAPI
Databases
Neo4j
pgvector
Qdrant
PostgreSQL
AI/ML
LLM
SGLang
vLLM
Knowledge Graph
AI Agents
DevOps
AWS
Azure
Bicep
GCP
Karpenter
KEDA
Kubernetes
OpenTelemetry
OpenTofu
Terraform
Vector
Cybersecurity
Least Privilege
Apply
$185k – $260k per year • Remote • Full-Time • 8+ years exp • Bachelor's Degree
DevOps
AWS
CI/CD
GCP
Kubernetes
GitHub
Cybersecurity
Clair
Dependabot
OWASP Top 10
OWASP ZAP
Snyk
Trivy
Apply
$105k – $252k per year • Remote • Full-Time • 18+ years exp • Bachelor's Degree
Python
Java
Java
Gradle
DevOps
Ansible
AWS
CI/CD
CloudFormation
Configuration Management
Docker
GitHub Actions
GitLab CI
Helm
Jenkins
Kubernetes
Platform Engineering
Terraform
GitHub
GitLab
Cybersecurity
Sonatype Nexus IQ
Management
Confluence
Jira
Apply
$54k – $175k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$35k – $113k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
Head of IT 13 days ago
$200k – $350k per year • In office • Full-Time • San Francisco
Python
AI/ML
Pre-training
Cybersecurity
Least Privilege
Management
Google Workspace
Slack
Apply
$225k – $550k per year • In office • Full-Time • San Francisco
C++
Go
Kotlin
Python
Rust
TypeScript
AI/ML
LLM
LLM Guardrails
Pre-training
DevOps
CI/CD
Cybersecurity
MITRE ATT&CK
Apply
$225k – $550k per year • In office • Full-Time • San Francisco
AI/ML
Pre-training
Apply
$225k – $550k per year • In office • Full-Time • San Francisco
AI/ML
Post-training
Pre-training
Apply
$200k – $550k per year • In office • Full-Time • San Francisco
AI/ML
Post-training
Pre-training
Apply
$180k – $210k per year • Equity • In office • Full-Time • San Francisco
Node JS
JavaScript
Databases
PostgreSQL
DevOps
PagerDuty
Web3
TRM Labs
Management
Slack
Apply
$252k – $335k per year • Remote/Hybrid • Full-Time • 8+ years exp • San Francisco
AI/ML
ChatGPT
Human-in-the-Loop
OpenAI
OpenAI Codex
DevOps
SLI/SLO/SLA
Apply
$223k – $424k per year (Estimated) • In office • Bachelor's Degree • San Francisco
AI/ML
AI Agents
LLM
Recommender Systems
Apply
$160k – $283k per year • Equity • In office • 5+ years exp • San Francisco
AI/ML
AI Agents
Apply
$185k – $385k per year • Remote/Hybrid • Full-Time • 5+ years exp • San Francisco
JavaScript
Python
Databases
MySQL
PostgreSQL
AI/ML
OpenAI
Frontend
React.js
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.