368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$191k – $388k per year (Estimated)
Location
In office (San Francisco, Tokyo)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Radical Numerics is an artificial intelligence research lab dedicated to developing general biological intelligence and generative genomics models. Headquartered in San Francisco, California, the company builds multimodal platforms capable of reading, writing, and engineering biological sequences across DNA, RNA, and proteins. Its technology aims to accelerate biopharmaceutical research, enhance early disease diagnostics, and establish robust biodefense capabilities.

About Us

Radical Numerics is an AI research lab building general biological intelligence. Our mission is to master the code of life, and our purpose is to reduce human suffering.

Our team created Evo, and started the field of generative genomics. Our work was featured on the cover of Science, and presented by our CEO on the main stage of TED2025. Evo was used to create the first AI gene therapy tool CRISPR-Cas9, and the first AI whole genome from scratch. Evo 2, featured in Nature, is the largest fully open source AI project across any domain.

Radical Numerics is bringing the rigor of distributed systems, model architecture, and numerics research to the challenges of biology. We’ve redesigned the foundation model training stack to turn the world’s raw scientific data (e.g. biological sequences, experiments, and physical processes), into intelligible, generative models that can expand and accelerate what humanity can understand, design, and cure.

The same generative breakthroughs that enable life-saving cures also lowers the barrier to creating engineered threats and AI-generated bioweapons. We believe these forces are inseparable. Radical Numerics was founded to develop both the power to design and the responsibility to defend.

About the Role

As a Member of Technical Staff, Distributed Systems at Radical Numerics, you will design, build, and operate the data center & cloud environment that powers our large-scale training and inference. You will deliver high-performance, reliable, and cost-efficient compute so our researchers can move fast at scale, turning frontier infrastructure into the foundation for the next generation of biological world models.

This role is ideal for someone who combines distributed systems, deep operational instincts and has an interest in modern machine learning. You should care about how every layer of the cluster affects research velocity: provisioning and capacity, scheduling and multi-tenancy, storage and lineage, communication overhead, observability, and the reliability of long-running jobs across thousands of accelerators.

What You'll Do

  • Operate and automate large GPU clusters. Own provisioning, imaging, and capacity planning across large distributed compute systems, with a focus on uptime, utilization, and cost efficiency.

  • Build a unified compute interface. Write software that abstracts cluster management and presents a single, ergonomic interface for training and inference, so researchers spend their time on science rather than infrastructure.

  • Extend scheduling and orchestration. Adapt systems like Kubernetes or Slurm for topology-aware placement, preemption, quotas, and fair-share multi-tenancy across competing workloads.

  • Maximize throughput and hardware efficiency. Profile and tune performance across the stack, including communication patterns, memory efficiency, custom kernels, compilation paths, and systems instrumentation, to ensure training compute is used effectively.

  • Improve reliability and recovery. Establish standards and mechanisms for robustness and error recovery, including monitoring, fault tolerance, checkpointing, and incident analysis for fast-moving research infrastructure.

  • Build reliable storage and artifact paths. Design durable paths for datasets, checkpoints, and logs, with clear retention and lineage that support reproducible, large-scale experimentation.

  • Collaborate across research and engineering. Partner closely with model researchers and training scientists to unblock large-scale runs, advise on parallelism and performance trade-offs, and design systems that support new scientific directions rather than constrain them.

What We're Looking For

  • Track record building distributed systems, or operating large-scale GPU clusters and container orchestration systems such as Kubernetes or Slurm.

  • Proficiency in building performant, maintainable software in at least one backend language (we use Python and Rust), with a focus on performance and reliability.

  • Strong systems background spanning Linux, networking, and infrastructure-as-code.

  • Strong understanding of modern deep learning frameworks and their systems internals (e.g., PyTorch, Triton, CUDA, C++).

  • Ability to debug complex, multi-layered systems involving distributed training, memory and performance regressions, and reliability issues in large codebases.

  • Comfort operating across the stack and owning projects end to end, with a bias toward initiative and execution.

  • Excellent written and verbal communication skills bridging technical and scientific domains.

Nice to Have

  • Familiarity with CUDA/NCCL and performance profiling for distributed training and inference.

  • Experience supporting large-scale distributed training for frontier or foundation models.

  • Contributions to open-source ML systems or infrastructure such as PyTorch, Torchtitan, or Megatron-LM.

  • Familiarity with ML runtimes, compilers, numerics, communication libraries, and custom kernel development.

  • Experience improving researcher productivity through infrastructure design, developer tooling, or workflow improvements.

  • Background in applied math, systems, computational biology, or related quantitative sciences.

Why Radical Numerics

  • Help build the computational foundation for multimodal biological world models aimed at rapid detection, response, and countermeasures across global health.

  • Work on systems problems at the frontier of distributed training, architecture, and numerics, in service of real biological applications.

  • Join a collaborative culture that values rigor, creativity, and cross-disciplinary partnership across AI labs, biotechs, hospital systems, and research institutes.

  • Competitive compensation, comprehensive benefits, and support for continual learning.

Radical Numerics is committed to equal employment opportunity and does not discriminate in any employment opportunities or practices based on an individual's race, color, creed, gender (including gender identity and gender expression), religion (all aspects of religious beliefs, observance or practice, including religious dress or grooming practices), marital status, registered domestic partner status, age, national origin or ancestry (including language use restrictions and possession of a driver’s license issued under California Vehicle Code section 12801.9), natural hair, physical or mental disability, political affiliation, medical condition (including cancer or a record or history of cancer, and genetic characteristics), sex (including pregnancy, childbirth, breastfeeding or related medical condition), genetic information, sexual orientation, military and veteran status or any other consideration made unlawful by federal, state, or local laws. It also prohibits unlawful discrimination based on the perception that anyone has any of those characteristics, or is associated with a person who has or is perceived as having any of those characteristics.

Radical Numerics participates in E-Verify and will provide the federal government with your Form I-9 information to confirm that you are authorized to work in the U.S.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$24k – $63k per year (Estimated) • In office • Full-Time • 6+ years exp • Master's Degree • India
Crystal
Groovy
JavaScript
Perl
Python
Ruby
SQL
TypeScript
Java
Java
Apache Tomcat
Gradle
Hibernate
Maven
Spring Boot
Spring MVC
Databases
Apache Kafka
Db2
Oracle
PostgreSQL
RabbitMQ
AI/ML
Fine-tuning
Frontend
Angular
JQuery
DevOps
Apache HTTP Server
AWS
Azure
CI/CD
Docker
GCP
Jenkins
Kubernetes
Rest API
Cybersecurity
Checkmarx
SonarQube
Apply
$96k – $163k per year • Equity • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Boston
C#
JavaScript
Python
TypeScript
DevOps
Azure AKS
CI/CD
Kubernetes
Rest API
Azure
Apply
$64k – $189k per year (Estimated) • Remote/Hybrid • Full-Time • 1+ year exp • Bachelor's Degree • Singapore
Python
SQL
Databases
Apache Kafka
AI/ML
Amazon SageMaker
Kubeflow
MLFlow
Spark
Vertex AI
DevOps
AWS
Azure
Azure DevOps
CI/CD
Docker
GCP
GitLab
GitLab CI
Jenkins
Kubernetes
Apply
$142k – $213k per year • Remote/Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • Jersey City
Java
Python
SQL
TypeScript
JavaScript
Java
Spring Boot
Python
Asyncio
FastAPI
AI/ML
AI Agents
Claude
Claude Code
Copilot
Cursor
Devin
Fine-tuning
Gemini
Google ADK
Hybrid Search
Knowledge Graph
LangChain
LangGraph
RAG
Frontend
Angular
DevOps
CI/CD
Docker
Kubernetes
Rest API
Apply
$20k – $52k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • India
Python
SQL
C#
TypeScript
JavaScript
C#
.NET
AI/ML
AI Agents
Frontend
Angular
Apply
$195k – $394k per year (Estimated) • In office • Full-Time • San Francisco
Python
AI/ML
Diffusion Models
Fine-tuning
JAX
PyTorch
Apply
$182k – $369k per year (Estimated) • In office • Full-Time • San Francisco
Python
AI/ML
LLM
TensorRT
TensorRT-LLM
Triton
vLLM
Apply
$193k – $391k per year (Estimated) • In office • Full-Time • San Francisco
Python
AI/ML
CUDA
CUDA Toolkit
DeepSpeed
LLM
Multimodal AI
PyTorch
SGLang
TensorRT
TensorRT-LLM
Triton
vLLM
Mixture of Experts
Apply
$200k – $406k per year (Estimated) • In office • Full-Time • San Francisco • Tokyo
Python
AI/ML
Multimodal AI
PyTorch
Apply
$193k – $391k per year (Estimated) • In office • Full-Time • San Francisco • Tokyo
Python
AI/ML
Multimodal AI
PyTorch
Pre-training
Apply
$223k – $424k per year (Estimated) • In office • Bachelor's Degree • San Francisco
AI/ML
AI Agents
LLM
Recommender Systems
Apply
$83k – $188k per year (Estimated) • In office • 2+ years exp • San Francisco
Python
AI/ML
AI Agents
LLM Guardrails
Model Context Protocol
DevOps
Terraform
Cybersecurity
Crowdstrike
GDPR
Least Privilege
Okta
SentinelOne
Management
Google Workspace
Slack
Apply
$171k – $273k per year • In office • Full-Time • 8+ years exp • PhD • San Francisco • Washington
AI/ML
A2A
Agentforce
AI Agents
Model Context Protocol
DevOps
AWS
GCP
Marketing
Salesforce
Apply
Security GRC Analyst 2 hours ago
$119k – $268k per year (Estimated) • Remote/Hybrid • 4+ years exp • Bachelor's Degree • San Francisco
AI/ML
Ignite
PyTorch
Cybersecurity
ISO 27001
NIST CSF
SOC 2
Apply
$173k – $260k per year • In office • Full-Time • PhD • San Francisco
JavaScript
Node JS
Python
Python
Celery
Django
Flask
Databases
RabbitMQ
Redis
AI/ML
Agentforce
AI Agents
DevOps
Akamai
AWS
CI/CD
Cloudflare
CloudFormation
Helm
Jenkins
Kubernetes
Spinnaker
Terraform
Marketing
Salesforce
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.