1,338,580open jobs
78,412companies
206,623added this week
Browse all
Salary
$180k – $360k per year
Location
Hybrid (San Francisco, New York, Seattle, United States, Toronto, Montreal, Canada)
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 8, 2026. First seen by Alion on Oct 5, 2026. Baseten scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Baseten is an American company founded in 2019 that runs machine learning models in production for companies that would rather not operate GPU infrastructure themselves. Its position is inference rather than training: it handles model packaging, autoscaling, cold start latency and multi-cloud capacity, which are the unglamorous problems that determine whether a model-powered product is fast and affordable enough to ship. Headquartered in San Francisco and backed at a multi-billion dollar valuation, it serves companies deploying open and custom models, and it competes with both the hyperscalers and the model providers' own hosted endpoints.

ABOUT BASETEN

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products.

THE ROLE

We're looking for inference performance engineers who want to make the world's most demanding AI workloads run faster and more efficiently. You'll work across the stack, from the inference engine and runtime through scheduling, serving, and routing. Along the way you'll apply techniques like prefill/decode disaggregation, speculative decoding, and KV-cache management. You'll reason from first principles about where time and memory go, find what's holding performance back, and close the gap. Your work directly impacts how fast our customers' models run and how efficiently we serve them. This role is ideal for someone who thrives in a fast-paced startup environment and is eager to make significant contributions to the exciting field of LLM inference.

EXAMPLE INITIATIVES

You'll get to work on these types of projects as an Inference Performance engineer:

RESPONSIBILITIES

  • Implement and productionize cutting-edge inference techniques, working deep in runtime internals. That includes quantization, speculative decoding, KV-cache reuse, chunked prefill, LoRA, guided generation for structured outputs, and custom scheduling and routing algorithms.

  • Profile and optimize inference end to end, from kernel launch overhead and memory layout up to request scheduling, prefill/decode disaggregation, and cache-aware routing. Run cross-layer investigations, such as tracing a tail-latency regression from request timing through routing and batching down to a kernel.

  • Turn performance into cost savings. Improve tokens per GPU-hour, raise utilization, and give customers and internal teams clear latency/throughput/cost tradeoffs.

  • Bring up and tune new model architectures on new hardware quickly, often in the same week they're released.

  • Build benchmarking frameworks that measure real-world performance across model architectures, batch sizes, sequence lengths, and hardware configurations.

  • Contribute upstream to open-source inference engines (vLLM, SGLang, TensorRT-LLM), and partner closely with model, infrastructure, and customer-facing teams to ship wins.

REQUIREMENTS

  • Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or related field.

  • Experience with one or more general-purpose programming languages, such as Python or C++.

  • Familiarity with LLM optimization techniques (e.g., quantization, speculative decoding, continuous batching).

  • Strong familiarity with ML libraries, especially PyTorch, TensorRT, or TensorRT-LLM.

  • Demonstrated interest and experience in LLMs.

  • Deep understanding of GPU architecture.

NICE TO HAVE

  • Proficiency in enhancing the performance of software systems, particularly in the context of large language models (LLMs)

  • Contributed to vLLM, SGLang, TensorRT-LLM, or another inference engine.

  • Worked on large-scale distributed serving: autoscaling, load balancing, multi-region or multi-cloud capacity.

  • Written or optimized GPU kernels (CUDA, Triton, CUTLASS, or similar)

  • Worked on quantization (FP8/FP4, AWQ, GPTQ) or speculative decoding in production.

  • Deep understanding of software engineering principles and a proven track record of developing and deploying AI/ML inference solutions.

BENEFITS

  • Competitive compensation, including meaningful equity

  • (U.S. only) 100% coverage of medical, dental, and vision insurance for employee and dependents

  • Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)

  • Paid parental leave

  • Fertility and family-building stipend through Carrot

  • (U.S. only) Company-facilitated 401(k)

  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.

At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.

We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,338,580 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
San Francisco
$185k – $246k per year • Hybrid • Full-Time • 15+ years exp • New York
AI/ML
AI Agents
LLM
LLM Guardrails
DevOps
CI/CD
AWS
Platform Engineering
Apply
$90k – $130k per year • In office • Full-Time • Bachelor's Degree • San Jose
JavaScript
TypeScript
C++
Frontend
Angular
DevOps
Wi-Fi
IoT
MQTT
Apply
$170k – $181k per year • Equity • Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Southlake
JavaScript
Java
Java
Spring Framework
Hibernate
Databases
MySQL
PostgreSQL
DevOps
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Management
Agile
Scrum
Apply
$85k – $92k per year • Equity • Hybrid • Full-Time • Bachelor's Degree • Southlake
JavaScript
TypeScript
SQL
Frontend
Angular
Management
Agile
Apply
$120k – $235k per year • In office • 4+ years exp • Bachelor's Degree • United States
Python
JavaScript
C#
C++
Databases
PostgreSQL
Azure Cosmos DB
Amazon DocumentDB
Azure SQL Database
Microsoft Fabric
DevOps
Azure
Linux
Analytics
Power BI
Azure Data Factory
Apply
≈ $102k – $304k per year (Estimated) • Hybrid • Internship • Master's Degree • Warren
Python
Python
pySpark
Databases
Databricks
AI/ML
Spark
TensorFlow
PyTorch
Machine Learning
DevOps
GCP
Azure
Apply
≈ $26k – $75k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Kamloops
Python
Java
C#
AI/ML
Model Context Protocol
AI Agents
NLP
LLM
Edge AI
Machine Learning
Management
Jira
Apply
$68k – $136k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Irving
Python
Management
Agile
Apply
up to $78k per year • Remote (Australia) • Freelance
Python
AI/ML
AI Agents
DevOps
GitHub
Apply
up to $78k per year • Remote (Canada) • Freelance
Python
AI/ML
AI Agents
DevOps
GitHub
Apply
$180k – $360k per year • Hybrid • Full-Time • Bachelor's Degree • San Francisco • Toronto • New York • Montreal • Seattle
AI/ML
Cursor
vLLM
Multimodal AI
Function Calling
SGLang
TensorRT
TensorRT-LLM
TGI
LLM
Structured Outputs
Baseten
KV Cache
Tool Use
Machine Learning
DevOps
CI/CD
Kubernetes
SLI/SLO/SLA
Management
Notion
Apply
$265k – $330k per year • Hybrid • Full-Time • San Francisco
AI/ML
Cursor
Baseten
Machine Learning
Management
Notion
Apply
$240k – $285k per year • Hybrid • Full-Time • 4+ years exp • San Francisco • Toronto • New York • Montreal • Seattle
Python
Go
AI/ML
Cursor
Spark
AI Agents
Baseten
Machine Learning
DevOps
Kubernetes
Management
Notion
Apply
$165k – $330k per year • Hybrid • Full-Time • San Francisco • New York • Toronto • Montreal • Seattle
Python
Go
Java
Java
Testcontainers
AI/ML
Cursor
Spark
Baseten
Machine Learning
DevOps
CI/CD
Docker
Kubernetes
Management
Notion
QA
JMeter
Gatling
Pytest
k6
Locust
Apply
$165k – $330k per year • Hybrid • Full-Time • San Francisco • New York • Toronto • Montreal • Seattle
Python
Go
AI/ML
Cursor
Spark
Baseten
Machine Learning
DevOps
CI/CD
GitOps
ArgoCD
Kubernetes
Progressive Delivery
Apply
≈ $132k – $222k per year (Estimated) • Remote (United States) • Full-Time • 8+ years exp • San Francisco
JavaScript
Java
TypeScript
SQL
Node JS
Java
Spring Boot
Databases
PostgreSQL
RabbitMQ
Apache Kafka
AI/ML
Copilot
Claude
Machine Learning
Frontend
Angular
React.js
DevOps
Rest API
CI/CD
Jenkins
AWS
Kubernetes
Configuration Management
QA
Swagger
Apply
$250k – $270k per year • Hybrid • Full-Time • 7+ years exp • San Francisco
Go
AI/ML
Model Context Protocol
AI Agents
Tool Use
DevOps
Rest API
Kong
Kubernetes
Apply
≈ $154k – $322k per year (Estimated) • In office • Full-Time • San Francisco
Python
Rust
C++
DevOps
WebRTC
eBPF
Robotics
Teleoperation
Apply
$90k – $110k per year • Hybrid • Full-Time • 3+ years exp • San Francisco • Los Angeles
Python
SQL
Databases
Snowflake
Amazon Redshift
DevOps
AWS
Amazon S3
Analytics
Tableau
Power BI
ETL/ELT
Alteryx
Looker
Domo
Microsoft Excel
Management
Asana
Smartsheet
Microsoft Teams
Apply
$100k – $300k per year • Equity 1–5% • Remote (United States) • Full-Time • 6+ years exp • San Francisco
Python
Rust
TypeScript
C++
Zig
Analytics
Microsoft Excel
Apply
See all jobs
This is one of many
1,338,580 more open roles from verified company boards, updated every day.