368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$156k – $342k per year (Estimated)
Location
In office (San Francisco)
Employment
Full-Time
Overview
Company
Impact
Profile match
COMPANY RESEARCH CLOUD Zyphra Cloud Login Two sides. Two sides.

Zyphra is an artificial intelligence company based in San Francisco, California.

The Role:

As a Research Engineer - AI Performance & Kernel Optimization, you will improve and optimize the performance of our large-scale language model training and inference stacks. You will work closely with our pretraining and inference teams to identify bottlenecks, design and implement highly optimized kernels, and push the limits of throughput, latency, and hardware utilization across a range of accelerator platforms. This role is suited for someone who enjoys deep systems work, cares about performance at every level of the stack, and is excited to translate low-level optimizations into meaningful gains for frontier-scale AI systems.

You’ll Work Across:

  • Kernel development and optimization for large-scale ML workloads, using any level of the stack from PTX/assembly to CUDA, HIP, Triton, or other GPU DSLs

  • Performance tuning for training and inference stacks across GPUs and other accelerators

  • Profiling and eliminating bottlenecks in memory movement, communication, scheduling, and compute utilization

  • Optimizing distributed training and inference systems for large MoE models, including large-scale model parallelism

  • Portability and optimization across non-NVIDIA hardware, with special interest in AMD hardware such as the MI300x and MI355x

  • Collaboration with research and infrastructure teams to turn systems improvements into real-world model training and inference gains

What We're Looking For / Requirements:

  • Strong engineering aptitude for building reliable, high-performance systems

  • Excellent low-level performance intuition and the ability to reason about hardware-software interactions

  • Are excited to rapidly learn new systems, tools, and hardware environments

  • Excellent communication and collaboration skills, with the ability to work effectively across research and engineering teams

  • Enjoy diving deep into the weeds and hunting down the last 10-20% of performance

Qualifications / Additional Skills:

  • Experience writing highly performant GPU kernels at any level of abstraction-PTX, CUDA, HIP, Triton, or other kernel DSLs

  • Experience optimizing ML workloads for large-scale training, ideally in language model pretraining or inference environments

  • Experience with non-NVIDIA accelerator hardware, such as AMD, AWS Trainium, Google TPU, Qualcomm, ARM, Intel, and custom ASICs

  • Strong understanding of distributed training systems and parallelism schemes, including data parallelism, tensor/model parallelism, pipeline parallelism, sharding, and communication/computation overlap

  • Experience with performance engineering in other demanding parallel computing environments such as HPC, quantitative finance, scientific computing, graphics, compilers, or numerical simulation

  • Strong systems intuition around memory hierarchy, bandwidth constraints, kernel fusion, launch overhead, communication overhead, and hardware utilization

  • Experience using profiling and debugging tools to drive performance improvements

  • Familiarity with infrastructure underlying large-scale training and inference, including collective communication libraries, and runtime performance analysis

  • Background in a highly technical field such as physics, mathematics, theoretical computer science, computer science, or electrical engineering

  • Any HPC experience is a strong plus

Why Work at Zyphra:

  • Our research methodology is grounded in methodical, step-by-step approaches to ambitious goals. Both deep research and engineering excellence are equally valued

  • We strongly value new and crazy ideas and are very willing to bet big on new ideas

  • We move as quickly as we can; we aim to minimize the bar to impact as low as possible

  • We all enjoy what we do and love discussing AI

Benefits and Perks:

  • Comprehensive medical, dental, vision, and FSA plans

  • Competitive compensation and 401(k) plan

  • Relocation and immigration support on a case-by-case basis

  • In-office snacks and meals provided

  • Unlimited PTO and company holidays

  • In-person team in San Francisco with a collaborative, high-energy environment

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$70k – $150k per year • In office • Full-Time • 5+ years exp • Calgary
Java
Node JS
Python
SQL
TypeScript
JavaScript
Java
Spring Boot
Databases
Apache Kafka
Chroma
OpenSearch
pgvector
Pinecone
PostgreSQL
AI/ML
AWS Bedrock
Copilot
Embeddings
LangChain
LangGraph
LLM
Prompt Engineering
RAG
AI Agents
Anthropic
Function Calling
Human-in-the-Loop
OpenAI
Structured Outputs
Frontend
Angular
React.js
DevOps
AWS
Azure
CI/CD
Docker
Kubernetes
Vector
Apply
Lead Data Engineer 3 hours ago
$140k – $231k per year • Remote/Hybrid • Full-Time • Bachelor's Degree • O'Fallon
Java
Python
SQL
Python
pySpark
Databases
Apache Kafka
Databricks
Delta Lake
AI/ML
Airflow
Hadoop
Spark
DevOps
AWS
Azure
CI/CD
GCP
Git
GitHub
Analytics
ETL/ELT
Apply
$115k – $184k per year • Remote/Hybrid • Full-Time • Bachelor's Degree • O'Fallon
JavaScript
SQL
TypeScript
Java
Java
Hibernate
Spring Boot
Spring Framework
Databases
Apache Kafka
PostgreSQL
Frontend
Angular
React.js
DevOps
AWS
Azure
CI/CD
Docker
GitHub
Kubernetes
Rest API
Apply
$73k – $186k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Toronto
Java
Python
SQL
Databases
Databricks
AI/ML
Spark
DevOps
AWS
Azure
Bitbucket
GCP
Git
Analytics
ETL/ELT
Apply
$88k – $224k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Waterloo
Java
Python
SQL
Databases
Databricks
AI/ML
Spark
DevOps
AWS
Azure
Bitbucket
GCP
Git
Analytics
ETL/ELT
Apply
$214k – $400k per year • Remote/Hybrid • 1+ year exp • Bachelor's Degree • San Francisco
C++
JavaScript
Python
SQL
TypeScript
AI/ML
AI Agents
Frontend
Angular
DevOps
CI/CD
Apply
$180k – $324k per year (Estimated) • In office • Full-Time • Master's Degree • San Francisco
AI/ML
AI Agents
Claude
Cursor
Fine-tuning
Multimodal AI
Apply
$158k – $345k per year (Estimated) • In office • Full-Time • San Francisco
Python
AI/ML
PyTorch
RAG
Reinforcement Learning
Pre-training
Apply
$161k – $352k per year (Estimated) • In office • Full-Time • San Francisco
Python
AI/ML
Fine-tuning
PyTorch
Reinforcement Learning
Synthetic Data
DPO
Post-training
SFT
Apply
$157k – $344k per year (Estimated) • In office • Full-Time • San Francisco
Python
AI/ML
PyTorch
Pre-training
Apply
$223k – $424k per year (Estimated) • In office • Bachelor's Degree • San Francisco
AI/ML
AI Agents
LLM
Recommender Systems
Apply
$83k – $188k per year (Estimated) • In office • 2+ years exp • San Francisco
Python
AI/ML
AI Agents
LLM Guardrails
Model Context Protocol
DevOps
Terraform
Cybersecurity
Crowdstrike
GDPR
Least Privilege
Okta
SentinelOne
Management
Google Workspace
Slack
Apply
$171k – $273k per year • In office • Full-Time • 8+ years exp • PhD • San Francisco • Washington
AI/ML
A2A
Agentforce
AI Agents
Model Context Protocol
DevOps
AWS
GCP
Marketing
Salesforce
Apply
Security GRC Analyst 2 hours ago
$119k – $268k per year (Estimated) • Remote/Hybrid • 4+ years exp • Bachelor's Degree • San Francisco
AI/ML
Ignite
PyTorch
Cybersecurity
ISO 27001
NIST CSF
SOC 2
Apply
$173k – $260k per year • In office • Full-Time • PhD • San Francisco
JavaScript
Node JS
Python
Python
Celery
Django
Flask
Databases
RabbitMQ
Redis
AI/ML
Agentforce
AI Agents
DevOps
Akamai
AWS
CI/CD
Cloudflare
CloudFormation
Helm
Jenkins
Kubernetes
Spinnaker
Terraform
Marketing
Salesforce
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.