600,779open jobs
30,816companies
86,462added this week
Browse all
Location
In office
Employment
Full-Time
Overview
Company
Impact
Profile match

About the Role

We’re looking for an AI Researcher focused on model distillation to help us push the frontier of efficient, high-performance models. You’ll work on turning large, expensive models into smaller, faster, and more deployable systems-while maintaining or improving quality.

This role is ideal for someone who enjoys publishing research, working close to real systems, and seeing their ideas move from papers → code → production.

What You’ll Work On

  • Design and evaluate model distillation techniques (teacher-student training, self-distillation, layer-wise distillation, representation matching, etc.)

  • Research tradeoffs between model size, latency, memory, and accuracy

  • Develop novel distillation approaches for:

    • Large language models

    • Long-context or specialized architectures

    • Inference-constrained environments

  • Run large-scale experiments and ablations; analyze results rigorously

  • Collaborate with engineers to productionize research outcomes

  • Write and submit research papers to top-tier venues (NeurIPS, ICML, ICLR, COLM, etc.)

  • Contribute to internal research notes, technical blogs, and open-source projects when appropriate

What We’re Looking For

Required

  • Strong background in machine learning research

  • Hands-on experience with model distillation or closely related topics (compression, pruning, quantization, representation learning)

  • Publication experience (conference or journal papers, workshop papers, or arXiv preprints)

  • Solid understanding of deep learning fundamentals (optimization, training dynamics, generalization)

  • Fluency in PyTorch (or equivalent) and research-grade experimentation

  • Ability to clearly communicate research ideas, results, and limitations

Nice to Have

  • Experience distilling large language models

  • Work on efficiency-focused research (latency, memory, throughput)

  • Experience with long-context models or non-Transformer architectures

  • Open-source contributions in ML or research tooling

  • Prior startup or applied research experience

Why Join Us

  • Real ownership over research direction at a Series A stage

  • Strong support for publishing and open research

  • Tight feedback loop between research and real-world deployment

  • Access to meaningful compute and production-scale problems

  • Small, highly technical team with deep ML and systems expertise

Example Backgrounds

  • ML researchers from academia transitioning to industry

  • Research engineers with published work in model efficiency

  • PhD / Post-doc graduates or industry researchers who still want to publish

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
600,779 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$124k – $196k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Santa Clara
Python
C
C++
Python
Dask
C
MPI
C++
PyTorch C++
AI/ML
LangGraph
LangChain
Spark
LlamaIndex
Model Context Protocol
CUDA Toolkit
Triton Inference Server
Fine-tuning
Reinforcement Learning
Quantization
Function Calling
AI Agents
TensorRT
TensorRT-LLM
PyTorch
CrewAI
LLM
RAPIDS
NVIDIA NIM
CUDA
Post-training
NVIDIA NeMo
OpenAI Agents SDK
A2A
NCCL
InfiniBand
NVLink
cuDNN
LLM Evaluation
LLM Guardrails
Multi-Agent Systems
Tool Use
DevOps
GCP
OpenShift
Azure
CI/CD
AWS
Kubernetes
GitHub
HPC
IoT
Matter
Apply
$184k – $288k per year • In office • Full-Time • 6+ years exp • Master's Degree • Santa Clara
Python
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
Multimodal AI
SGLang
PyTorch
LLM
CUDA
Triton
NCCL
CUTLASS
Management
Agile
Apply
$30k – $79k per year (Estimated) • In office • Full-Time • 4+ years exp • Master's Degree • Shanghai
C++
C++
PyTorch C++
AI/ML
CUDA Toolkit
TensorRT
PyTorch
LLM
CUDA
cuDNN
Apply
$100k – $167k per year • In office • Full-Time • Bachelor's Degree • Santa Clara
Python
AI/ML
Pandas
NumPy
PyTorch
DevOps
HPC
Apply
$245k – $279k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • McLean • San Francisco • New York
AI/ML
NeMo Guardrails
PyTorch
LLM
Hugging Face
LLM Guardrails
DevOps
GCP
Azure
AWS
Apply
$180k – $240k per year • Remote • Full-Time • Atlanta • Los Angeles • Tampa • Miami • Orlando
AI/ML
Claude
ChatGPT
Marketing
HubSpot
Apply
$150k – $190k per year • Remote • Full-Time • Atlanta • Tampa • Miami • Orlando • Ottawa
AI/ML
Claude
ChatGPT
Marketing
HubSpot
Apply
Chief of Staff 1 month ago
Remote/Hybrid • Full-Time • San Francisco
Management
Discord
Apply
Content Marketer 3 months ago
$42k – $56k per year • Remote • Full-Time
Apply
$42k – $112k per year (Estimated) • Remote • Full-Time • 1+ year exp • Prague • Lisbon • London • Milan • Munich
AI/ML
OpenAI
Management
n8n
Marketing
HubSpot
Apply
See all jobs
This is one of many
600,779 more open roles from verified company boards, updated every day.