368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$225k – $550k per year
Location
In office (San Francisco)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Magic is a San Francisco research company founded in 2022 that trains frontier models for software engineering. It focuses on extremely long context windows so a model can hold entire codebases and their history in working memory while making changes. The company raised large rounds from investors including Alphabet's CapitalG and Nat Friedman and builds its own training infrastructure.

Magic’s mission is to build safe AGI that accelerates humanity’s progress on the world’s most important problems. We believe the most promising path to safe AGI lies in automating research and code generation to improve models and solve alignment more reliably than humans can alone. Our approach combines frontier-scale pre-training, domain-specific RL, ultra-long context, and inference-time compute to achieve this goal.

About the role

As a Software Engineer on the Pre-training Systems team, you will design and operate the distributed infrastructure that trains Magic’s long-context models at scale.

This role focuses on large-scale model training across massive GPU clusters. You will work at the boundary between deep learning and distributed systems, ensuring that training runs are performant, reliable, and reproducible under extreme scale.

Magic’s long-context models create non-trivial systems challenges: sustained memory pressure, communication overhead across thousands of devices, long-running jobs that must survive failures, and efficient sequence packing under hardware constraints. You will own the systems that make large-scale pre-training stable and fast.

What you’ll work on

  • Scale distributed training across large GPU clusters (data, tensor, pipeline parallelism)

  • Optimize communication patterns and gradient synchronization

  • Improve checkpointing, fault tolerance, and job recovery systems

  • Profile and eliminate performance bottlenecks across compute, networking, and storage

  • Improve experiment reproducibility and orchestration workflows

  • Increase hardware utilization and training throughput

  • Collaborate with Kernels and Research to align model architecture with systems realities

What we’re looking for

  • Strong software engineering and distributed systems fundamentals

  • Experience training large models in multi-node GPU environments

  • Deep understanding of parallelism strategies and performance trade-offs

  • Experience debugging cross-layer issues in production ML systems

  • Strong ownership mindset and ability to operate critical infrastructure

  • Track record of improving performance or reliability of large-scale systems

Our culture

  • Integrity. Words and actions should be aligned

  • Hands-on. At Magic, everyone is building

  • Teamwork. We move as one team, not N individuals

  • Focus. Safely deploy AGI. Everything else is noise

  • Quality. Magic should feel like magic

Magic strives to be the place where high-potential individuals can do their best work. We value quick learning and grit just as much as skill and experience.

Compensation, benefits, and perks (US):

  • Annual salary range: $225K - $550K

  • Equity is a significant part of total compensation, in addition to salary

  • 401(k) plan with 6% salary matching

  • Generous health, dental and vision insurance for you and your dependents

  • Unlimited paid time off

  • Visa sponsorship and relocation stipend to bring you to SF, if possible

  • A small, fast-paced, highly focused team

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
Research Scientist 3 days ago
$165k – $220k per year • Remote/Hybrid • Full-Time • 5+ years exp • PhD • New York • Boston
JavaScript
Python
TypeScript
AI/ML
Post-training
Pre-training
AI Agents
Apply
$41k – $64k per year • Remote • 3+ years exp
Python
AI/ML
LLM
NLP
Transformers
Megatron-LM
Pre-training
DevOps
HPC
Apply
$200k – $420k per year • In office • Bachelor's Degree • Palo Alto
Python
Rust
TypeScript
JavaScript
AI/ML
LLM
Pre-training
Structured Outputs
Function Calling
Frontend
React.js
Mobile
React Native
DevOps
CI/CD
Kubernetes
Terraform
Apply
Remote • Full-Time
Python
AI/ML
LLM
Prompt Engineering
Synthetic Data
Tokenization
Post-training
Pre-training
AI Agents
Apply
$180k – $450k per year • In office • Full-Time • San Jose
Python
AI/ML
Fine-tuning
Knowledge Distillation
LLM
Multimodal AI
PyTorch
Reinforcement Learning
Synthetic Data
Post-training
Pre-training
AI Agents
Function Calling
Robotics
Reinforcement Learning
Apply
Head of IT 13 days ago
$200k – $350k per year • In office • Full-Time • San Francisco
Python
AI/ML
Pre-training
Cybersecurity
Least Privilege
Management
Google Workspace
Slack
Apply
$225k – $550k per year • In office • Full-Time • San Francisco
C++
Go
Kotlin
Python
Rust
TypeScript
AI/ML
LLM
LLM Guardrails
Pre-training
DevOps
CI/CD
Cybersecurity
MITRE ATT&CK
Apply
$225k – $550k per year • In office • Full-Time • San Francisco
AI/ML
Post-training
Pre-training
Apply
$200k – $550k per year • In office • Full-Time • San Francisco
AI/ML
Post-training
Pre-training
Apply
$200k – $550k per year • In office • Full-Time • San Francisco
AI/ML
Pre-training
DevOps
AWS
Azure
GCP
Kubernetes
Terraform
Apply
$180k – $210k per year • Equity • In office • Full-Time • San Francisco
Node JS
JavaScript
Databases
PostgreSQL
DevOps
PagerDuty
Web3
TRM Labs
Management
Slack
Apply
$252k – $335k per year • Remote/Hybrid • Full-Time • 8+ years exp • San Francisco
AI/ML
ChatGPT
Human-in-the-Loop
OpenAI
OpenAI Codex
DevOps
SLI/SLO/SLA
Apply
$223k – $424k per year (Estimated) • In office • Bachelor's Degree • San Francisco
AI/ML
AI Agents
LLM
Recommender Systems
Apply
$160k – $283k per year • Equity • In office • 5+ years exp • San Francisco
AI/ML
AI Agents
Apply
$185k – $385k per year • Remote/Hybrid • Full-Time • 5+ years exp • San Francisco
JavaScript
Python
Databases
MySQL
PostgreSQL
AI/ML
OpenAI
Frontend
React.js
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.