611,334open jobs
33,302companies
86,098added this week
Browse all
Salary
$165k – $225k per year
Location
In office (San Francisco)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match

Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.

About the role

As a Pre-training Research Engineer, you’ll focus on model implementation, pertaining and scaling, and improving the quality of our byte-native and multimodal foundation models. You’ll build and iterate quickly on research ideas, contribute production-grade training code and infrastructure, and help deliver high-quality base models that can serve real-world use cases at scale.

Key Responsibilities

Pre-training & Scaling

  • Train large byte-native and multimodal foundation models across massive, heterogeneous corpora.

  • Implement and evaluate new model architectures, training objectives, and optimization methods.

  • Develop stable pre-training recipes and run scaling experiments for novel architectures.

  • Conduct ablations and analyze training dynamics, model behavior, and base-model quality.

  • Work with data and distributed training engineers to improve training efficiency, reliability, and scalability.

Must-Haves

  • 5+ years of experience in machine learning research or engineering, with a proven track record of developing and pre-training large language or multimodal foundation models.

  • Software Engineering: Strong general software engineering skills, with the ability to write robust and performant training code.

  • ML Foundations: Solid understanding of deep learning fundamentals and modern pre-training methods and literature.

  • Research and Experimentation: Ability to quickly implement research ideas and evaluate them using clear baselines, ablations, metrics, and analysis.

  • GPU and Distributed Training: Hands-on experience running training workloads in GPU-based environments, with familiarity with distributed training.

  • Education: MS in Computer Science, Machine Learning, Artificial Intelligence, Mathematics, or a related field.

Nice-to-Haves

  • PhD in Computer Science, Machine Learning, Artificial Intelligence, Mathematics, or a related field.

  • JAX Ecosystem: Extensive experience with the JAX, Flax, and XLA stack.

  • Large-Scale Distributed Training: Experience with multi-node pre-training using systems such as FSDP, ZeRO, or Megatron.

  • Training Recipes and Scaling: Experience developing training recipes, ablations, or scaling experiments.

  • Monitoring and Reproducibility: Experience owning end-to-end training and evaluation pipelines with monitoring and reproducibility.

Education

  • MS or PhD in Computer Science, Machine Learning, Artificial Intelligence, Mathematics, or a related field.

Benefits include

  • Medical, dental, and vision insurance

  • 401k plan

  • Daily lunch, snacks, and beverages

  • Flexible time off

  • Competitive salary and equity

Equal opportunity

Sciforium is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
611,334 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$93k – $139k per year • In office • Full-Time • 3+ years exp • Berlin
Python
JavaScript
TypeScript
SQL
Node JS
Node JS
Fastify
Prisma
AI/ML
Prefect
Multimodal AI
AI Agents
Cohere SDK
Langfuse
Gemini
LLM
OpenAI
OCR
Semantic Search
DevOps
Datadog
Azure
AWS
Vector
Amazon S3
Analytics
ETL/ELT
Apply
$158k – $303k per year (Estimated) • In office • Full-Time • Master's Degree • Sunnyvale
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
JAX
Multimodal AI
Computer Vision
TensorRT
TensorFlow
PyTorch
Self-Supervised Learning
Human-in-the-Loop
Edge AI
Vision-Language-Action
Embodied AI
Robotics
Sim-to-Real
Apply
$100k – $150k per year • Remote • 6+ years exp • Master's Degree
Python
AI/ML
Fine-tuning
JAX
Multimodal AI
AI Agents
PyTorch
RAG
Apply
$100k – $150k per year • Remote • 6+ years exp • Master's Degree
Python
AI/ML
LoRA
Fine-tuning
RLHF
Reinforcement Learning
Multimodal AI
Knowledge Distillation
PEFT
QLoRA
PyTorch
LLM
Synthetic Data
DPO
FSDP
Model Distillation
Apply
$90k – $115k per year • Remote • 6+ years exp • Master's Degree
Python
AI/ML
Fine-tuning
RLHF
Reinforcement Learning
Multimodal AI
Knowledge Distillation
PyTorch
LLM
Synthetic Data
DPO
FSDP
Model Distillation
Apply
GPU Kernel Engineer 4 days ago
$190k – $250k per year • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco
Python
C++
Assembly
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
JAX
Multimodal AI
TensorRT
PyTorch
LLM
CUDA
Triton
TPU
ROCm
Edge AI
XLA
Apply
$155k – $200k per year • In office • Full-Time • 2+ years exp • Master's Degree • San Francisco
Python
AI/ML
vLLM
JAX
Multimodal AI
SGLang
TensorRT
TensorRT-LLM
Transformers
TensorFlow
PyTorch
Hugging Face
ROCm
DevOps
CI/CD
Apply
$190k – $250k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco
Python
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
JAX
Multimodal AI
TensorRT
PyTorch
CUDA
TPU
NCCL
ROCm
XLA
Apply
$230k – $300k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco
Python
C++
AI/ML
vLLM
CUDA Toolkit
Multimodal AI
SGLang
LLM
Ray
CUDA
ROCm
KV Cache
DevOps
Kubernetes
HPC
Apply
$165k – $210k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • San Francisco
Python
TypeScript
C++
Python
FastAPI
AI/ML
vLLM
Multimodal AI
SGLang
LLM
DevOps
WebSockets
Management
Stripe
Apply
$120k – $243k per year (Estimated) • In office • Full-Time • 9+ years exp • Bachelor's Degree • San Francisco
Python
SQL
Analytics
Tableau
Metabase
Looker
Apply
$109k – $230k per year (Estimated) • In office • Full-Time • 3+ years exp • San Francisco
Design
Figma
Sketch
Adobe XD
Apply
$150k – $250k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • San Francisco
TypeScript
Databases
PostgreSQL
AI/ML
AI Agents
DevOps
AWS
Kubernetes
Apply
$100k – $180k per year • In office • Full-Time • San Francisco
Python
TypeScript
AI/ML
Computer Vision
Frontend
Tailwind CSS
Apply
$150k – $200k per year • In office • Full-Time • 8+ years exp • San Francisco
DevOps
Azure
Cybersecurity
Okta
Management
Slack
ServiceNow
Microsoft Teams
Marketing
LinkedIn
Apply
See all jobs
This is one of many
611,334 more open roles from verified company boards, updated every day.