368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$180k – $360k per year
Location
Remote/Hybrid (San Francisco, New York, United States, Toronto, Montreal, Canada)
Employment
Full-Time
Overview
Company
Impact
Profile match
Baseten is an AI infrastructure platform designed to help developers and machine learning teams deploy, serve, and scale open-source and custom AI models. The platform provides performant, low-latency inference infrastructure alongside developer tools like Truss, an open-source model packaging framework. Headquartered in San Francisco, California, Baseten enables companies to run state-of-the-art models in production seamlessly without managing underlying cloud infrastructure.

ABOUT BASETEN

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products.

THE ROLE

Are you passionate about advancing the application of artificial intelligence? We are looking for a Software Engineer focused on ML performance to join our dynamic team. This role is ideal for someone who thrives in a fast-paced startup environment and is eager to make significant contributions to the exciting field of LLM Inference. If you are a backend engineer who thrives on making things faster and is excited about open-source ML models, we look forward to your application.

EXAMPLE INITIATIVES

You'll get to work on these types of projects as part of our Model Performance team:

RESPONSIBILITIES

  • Implement, refine, and productionize cutting-edge techniques (quantization, speculative decoding, kv cache reuse, chunked prefill and LoRA) for ML model inference and infrastructure.

  • Deep dive into underlying codebases of TensorRT, PyTorch, TensorRT-LLM, vllm, sglang, CUDA, and other libraries to debug ML performance issues.

  • Apply and scale optimization techniques across a wide range of ML models, particularly large language models.

  • Collaborate with a diverse team to design and implement innovative solutions.

  • Own projects from idea to production.

REQUIREMENTS

  • Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or related field.

  • Experience with one or more general-purpose programming languages, such as Python or C++.

  • Familiarity with LLM optimization techniques (e.g., quantization, speculative decoding, continuous batching).

  • Strong familiarity with ML libraries, especially PyTorch, TensorRT, or TensorRT-LLM.

  • Demonstrated interest and experience in LLM’s.

  • Deep understanding of GPU architecture.

  • Bonus:

    • Proficiency in enhancing the performance of software systems, particularly in the context of large language models (LLMs).

    • Experience with CUDA or similar technologies.

    • Deep understanding of software engineering principles and a proven track record of developing and deploying AI/ML inference solutions.

    • Experience with Docker and Kubernetes.

BENEFITS

  • Competitive compensation, including meaningful equity.

  • 100% coverage of medical, dental, and vision insurance for employee and dependents

  • Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)

  • Paid parental leave

  • Fertility and family-building stipend through Carrot

  • Company-facilitated 401(k)

  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.

At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.

We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$145k – $218k per year • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Mississauga
Python
Python
FastAPI
Databases
FAISS
pgvector
Pinecone
Weaviate
PostgreSQL
AI/ML
AI Agents
Anthropic
Claude
Embeddings
Fine-tuning
Gemini
Hallucination
Hugging Face
Human-in-the-Loop
Hybrid Search
LangChain
LangGraph
Llama
LlamaIndex
LLM
LLM Guardrails
OpenAI
Prompt Engineering
PyTorch
RAG
TensorFlow
Mobile
Clean Architecture
DevOps
Docker
Incident Management
Vector
Apply
$115k – $228k per year (Estimated) • Remote/Hybrid • 6+ years exp • Bachelor's Degree • Scottsdale
MATLAB
SystemVerilog
TCL Scripting
Verilog
VHDL
MATLAB
Simulink
AI/ML
Claude
Claude Code
Copilot
Cursor
Quantization
DevOps
Git
GitHub
Chips/EDA
HDL Coder
Xilinx Vivado
Apply
$110k – $219k per year (Estimated) • Equity • Remote/Hybrid • 6+ years exp • Bachelor's Degree • Boston
MATLAB
SystemVerilog
TCL Scripting
Verilog
VHDL
MATLAB
Simulink
AI/ML
Claude
Claude Code
Copilot
Cursor
Quantization
DevOps
Git
GitHub
Chips/EDA
HDL Coder
Xilinx Vivado
Apply
$18k – $47k per year (Estimated) • In office • Bachelor's Degree • Moscow
C++
MATLAB
Python
Databases
ClickHouse
DevOps
Grafana
Robotics
ROS2
RViz
ROS
SpaceTech
QGIS
Apply
$26k – $68k per year (Estimated) • In office • Full-Time • 14+ years exp • Pune
C++
Python
DevOps
CI/CD
Git
Apply
$165k – $330k per year • Remote/Hybrid • Full-Time • San Francisco • Toronto • New York • Montreal
AI/ML
Cursor
DeepSeek
LLM
SGLang
TensorRT
TensorRT-LLM
vLLM
Management
Notion
Apply
$190k – $240k per year • Remote/Hybrid • Full-Time • San Francisco • New York
JavaScript
TypeScript
AI/ML
Cursor
Frontend
Next.js
React Three Fiber
React.js
Three.JS
Design
Figma
Apply
Product Designer 5 days ago
$225k – $290k per year • Remote/Hybrid • Full-Time • San Francisco • New York
JavaScript
TypeScript
AI/ML
Cursor
Frontend
React.js
DevOps
Git
Apply
AI Engineer 12 days ago
$220k – $260k per year • Remote/Hybrid • Full-Time • 5+ years exp • San Francisco
Python
AI/ML
Claude
Claude Code
Cursor
Fine-tuning
LangChain
LLM
LoRA
Reinforcement Learning
Spark
Synthetic Data
PEFT
LLM Guardrails
OpenAI Codex
Post-training
SFT
AI Agents
Model Context Protocol
Apply
$170k – $220k per year • Remote/Hybrid • Full-Time • 8+ years exp • San Francisco • New York
Go
AI/ML
Cursor
LLM
Post-training
Apply
$180k – $210k per year • Equity • In office • Full-Time • San Francisco
Node JS
JavaScript
Databases
PostgreSQL
DevOps
PagerDuty
Web3
TRM Labs
Management
Slack
Apply
$252k – $335k per year • Remote/Hybrid • Full-Time • 8+ years exp • San Francisco
AI/ML
ChatGPT
Human-in-the-Loop
OpenAI
OpenAI Codex
DevOps
SLI/SLO/SLA
Apply
$223k – $424k per year (Estimated) • In office • Bachelor's Degree • San Francisco
AI/ML
AI Agents
LLM
Recommender Systems
Apply
$160k – $283k per year • Equity • In office • 5+ years exp • San Francisco
AI/ML
AI Agents
Apply
$185k – $385k per year • Remote/Hybrid • Full-Time • 5+ years exp • San Francisco
JavaScript
Python
Databases
MySQL
PostgreSQL
AI/ML
OpenAI
Frontend
React.js
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.