368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$200k – $500k per year
Location
Remote/Hybrid (San Francisco, United States)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Tzafon is advancing machine intelligence. Our team develops models and systems for autonomous agents interacting with the real world.

Tzafon is a foundation model lab building scalable compute systems and advancing machine intelligence, with offices in San Francisco, Zurich & Tel Aviv. We’ve raised over $12m in funding to advance our mission of expanding the frontiers of machine intelligence.

We're a team of engineers and scientists with deep backgrounds in ML infrastructure & research. Founded by IOI and IMO medalists, PhDs, and alumni from leading tech companies, such as Google Deepmind, Character, and NVIDIA, we train models and build infrastructure for swarms of agents to automate work across real-world environments.

You'll work between our product and post-training teams to ship Large Action Models that actually work. Build evals, benchmarks, and fine-tuning pipelines. Define what good model behavior means and make it happen at scale.

What you'll do

  • Design and execute large scale training runs on our clusters

  • Build and optimize distributed training infrastructure across massive multi-node systems

  • Implement post-training pipelines at scale

  • Develop data pipelines that process and filter trillions of tokens for pre-training

  • Research and implement architectural improvements, scaling laws, and training optimizations

  • Debug training instabilities, loss spikes, and convergence issues in long-running jobs

  • Build tooling for cluster utilization, fault tolerance, and checkpoint management

  • Write custom CUDA/Triton kernels to optimize critical training operations (attention, normalization, activations)

  • Collaborate on research that advances the state of the art in foundation model training

We're looking for

  • Deep experience pre-training or post-training foundation models on large clusters

  • Expert-level at Python and ML frameworks (PyTorch, JAX, Torchtitan)

  • Strong systems skills: distributed training, FSDP/ZeRO, tensor parallelism, pipeline parallelism

  • Experience writing performant CUDA or Triton kernels for ML workloads

  • Track record of running stable multi-week training jobs and debugging distributed training failures

  • Understanding of cluster scheduling, networking bottlenecks, and GPU/TPU performance optimization

Preferred Experience

  • Trained foundation models at major AI labs (OpenAI, Anthropic, Google DeepMind, Meta, xAI, etc.)

  • Worked on large scale RL runs

  • Optimized critical training kernels (FlashAttention, fused optimizers, custom kernels)

  • Published research at top ML conferences (NeurIPS, ICML, ICLR)

  • Contributions to open source ML infrastructure (PyTorch, JAX, vLLM, etc.)

  • Experience with training data pipelines, data quality research, or synthetic data generation

Life at Tzafon

  • Full medical, dental, and vision coverage, plus 401(k) in the us

  • Office in SF, Zurich, and Tel Aviv

  • Early-stage equity in a future-defining company

Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.

Compensation starts at $200k-$500k + equity package, depending on experience & location.

We also offer a referral bonus of $5k for referral of successful hires (send to [email protected]).

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$122k – $200k per year • In office • Full-Time • 10+ years exp • Jersey City • Charlotte
Java
Node JS
Python
SQL
JavaScript
Java
Hibernate
Spring Framework
Spring MVC
Databases
Apache Kafka
DynamoDB
Oracle
RabbitMQ
AI/ML
Ray
DevOps
Amazon CloudWatch
Amazon EC2
Amazon ECS
Amazon EKS
Amazon S3
Ansible
API Gateway
AWS
AWS Lambda
Azure
CI/CD
CloudFormation
GCP
Git
Grafana
IAM
Jenkins
Prometheus
Splunk
Terraform
Kubernetes
Apply
$111k – $200k per year (Estimated) • In office • Full-Time • 1+ year exp • Plano
JavaScript
Python
C#
C#
.NET
DevOps
Azure
CI/CD
Apply
$215k – $393k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • High School Diploma • San Francisco • Mountain View
AI/ML
JAX
TensorFlow
TPU
Apply
$71k – $112k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Austria
C#
C++
Python
Visual Basic
Apply
$106k – $149k per year • Remote/Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • Overland Park
Python
SQL
Databases
Amazon Redshift
Analytics
Power BI
Tableau
Marketing
Salesforce
Apply
$150k – $250k per year • Equity • In office • Full-Time • San Francisco
Rust
Databases
Apache Kafka
ClickHouse
NATS
PostgreSQL
Redis
DevOps
gRPC
Kubernetes
OpenTelemetry
Prometheus
WebSockets
Amazon S3
Cybersecurity
SOC 2
Apply
$150k – $400k per year • Equity • In office • Full-Time • San Francisco • Zurich • Tel Aviv
Python
AI/ML
Fine-tuning
LLM
MLFlow
PyTorch
RLHF
DPO
Post-training
Weights & Biases
AI Agents
Apply
$223k – $424k per year (Estimated) • In office • Bachelor's Degree • San Francisco
AI/ML
AI Agents
LLM
Recommender Systems
Apply
$83k – $188k per year (Estimated) • In office • 2+ years exp • San Francisco
Python
AI/ML
AI Agents
LLM Guardrails
Model Context Protocol
DevOps
Terraform
Cybersecurity
Crowdstrike
GDPR
Least Privilege
Okta
SentinelOne
Management
Google Workspace
Slack
Apply
$171k – $273k per year • In office • Full-Time • 8+ years exp • PhD • San Francisco • Washington
AI/ML
A2A
Agentforce
AI Agents
Model Context Protocol
DevOps
AWS
GCP
Marketing
Salesforce
Apply
Security GRC Analyst 2 hours ago
$119k – $268k per year (Estimated) • Remote/Hybrid • 4+ years exp • Bachelor's Degree • San Francisco
AI/ML
Ignite
PyTorch
Cybersecurity
ISO 27001
NIST CSF
SOC 2
Apply
$173k – $260k per year • In office • Full-Time • PhD • San Francisco
JavaScript
Node JS
Python
Python
Celery
Django
Flask
Databases
RabbitMQ
Redis
AI/ML
Agentforce
AI Agents
DevOps
Akamai
AWS
CI/CD
Cloudflare
CloudFormation
Helm
Jenkins
Kubernetes
Spinnaker
Terraform
Marketing
Salesforce
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.