368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$220k – $320k per year
Location
In office (San Francisco)
Seniority
Senior · 2+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Models. Agents. GPUs. Whatever you need — we've got it. Ghost agent VMs, Maestro model routing, Engine wholesale GPUs, and Academy, on one platform.

Help us make inference blazingly fast. If you love squeezing every last drop of performance out of GPUs, diving deep into CUDA kernels, and turning optimization techniques into production systems, we'd love to meet you.

AboutInference.net

Inference.net trains and hosts specialized language models for companies that need frontier-quality AI at a fraction of the cost. The models we train match GPT-5 accuracy but are smaller, faster, and up to 90% cheaper. Our platform handles everything end-to-end: distillation, training, evaluation, and planet-scale hosting.

We are a well-funded ten-person team of engineers who work in-person in downtown San Francisco on difficult, high-impact engineering problems. Everyone on the team has been writing code for over 10 years, and has founded and run their own software companies. We are high-agency, adaptable, and collaborative. We value creativity alongside technical prowess and humility. We work hard, and deeply enjoy the work that we do. Most of us are in the office 4 days a week in SF; hybrid works for Bay Area candidates.

About the Role

You will be responsible for making our inference stack as fast and efficient as possible. Your work spans from implementing known optimization techniques to experimenting with novel approaches, always with the goal of serving models faster and cheaper at scale.

Your north star is inference performance: latency, throughput, cost efficiency, and how quickly we can bring new model architectures into production. You'll work across the full inference stack-from CUDA kernels to serving frameworks-to find and eliminate bottlenecks. This role reports directly to the founding team. You'll have the autonomy, a large compute budget, and technical support to push the limits of what's possible in model serving.

Key Responsibilities

  • Implement and productionize optimization techniques including quantization, speculative decoding, KV cache optimization, continuous batching, and LoRA serving

  • Deep dive into inference frameworks (vLLM, SGLang, TensorRT-LLM) and underlying libraries to debug and improve performance

  • Profile and optimize CUDA kernels and GPU utilization across our serving infrastructure

  • Add support for new model architectures, ensuring they meet our performance standards before going to production

  • Experiment with novel inference techniques and bring successful approaches into production

  • Build tooling and benchmarks to measure and track inference performance across our fleet

  • Collaborate with applied ML engineers to ensure trained models can be served efficiently

Requirements

  • 2+ years of experience in ML systems, inference optimization, or GPU programming

  • Strong proficiency in Python and familiarity with C++

  • Hands-on experience with LLM inference frameworks (vLLM, SGLang, TensorRT-LLM, or similar)

  • Deep understanding of GPU architecture and experience profiling GPU workloads

  • Familiarity with LLM optimization techniques (quantization, speculative decoding, continuous batching, KV cache management)

  • Experience with PyTorch and understanding of how models execute on hardware

  • Track record of measurably improving system performance

Nice-to-Have

  • Experience with CUDA programming

  • Familiarity with serving non-LLM models (TTS, vision, embeddings)

  • Experience with distributed inference and multi-GPU serving

  • Contributions to open-source inference frameworks

  • Experience with Docker and Kubernetes

You don't need to tick every box. Curiosity and the ability to learn quickly matter more.

Compensation

We offer competitive compensation, equity in a high-growth startup, and comprehensive benefits. The base salary range for this role is $220,000 - $320,000, plus equity and benefits, depending on experience.

Equal Opportunity

Inference.net is an equal opportunity employer. We welcome applicants from all backgrounds and don't discriminate based on race, color, religion, gender, sexual orientation, national origin, genetics, disability, age, or veteran status.

If you're excited about making AI inference faster for everyone, we'd love to hear from you. Please send your resume and GitHub to [email protected] and/or apply here on Ashby.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
Data Scientist 1 day ago
$25k – $56k per year (Estimated) • Remote • Bachelor's Degree • Moscow
C++
Python
SQL
AI/ML
Computer Vision
CUDA
CUDA Toolkit
TensorRT
DevOps
Docker
Git
Kubernetes
Apply
AI Engineer 1 day ago
$25k – $103k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Gurgaon
Python
SQL
Databases
Databricks
Microsoft Fabric
AI/ML
AI Agents
Embeddings
Gemini
Hallucination
LangChain
LangGraph
LLM
Multimodal AI
Prompt Engineering
PyTorch
RAG
Semantic Search
Spark
TensorFlow
Hugging Face
LLM Guardrails
LLMOps
OpenAI
Semantic Search
DevOps
AWS
Azure
CI/CD
Apply
$77k – $215k per year (Estimated) • In office • Contractor • 3+ years exp • Singapore
JavaScript
Python
TypeScript
Python
Django
FastAPI
Databases
PostgreSQL
Redis
AI/ML
AI Agents
LangGraph
LLM
Prompt Engineering
LangChain
Edge AI
Frontend
Angular
React.js
Vue.js
Apply
$71k – $170k per year (Estimated) • In office • Full-Time • Netanya
Python
TypeScript
AI/ML
Accelerate
Fine-tuning
LangChain
LLM
NLP
Prompt Engineering
PyTorch
RAG
Edge AI
Hugging Face
AI Agents
DevOps
AWS
Azure
Docker
GCP
Kubernetes
Apply
$16k – $36k per year (Estimated) • Remote • Full-Time • Moscow
AI/ML
ChatGPT
Claude
Claude Code
Cursor
LLM
Model Context Protocol
Apply
$250k – $350k per year • In office • Full-Time • 3+ years exp • San Francisco
C#
C#
.NET
AI/ML
DeepSpeed
Knowledge Distillation
LLM
Multimodal AI
PyTorch
Reinforcement Learning
RLHF
Transformers
TRL
DPO
GPT-5
Hugging Face
Megatron-LM
Post-training
SFT
DevOps
GitHub
Apply
$220k – $320k per year • In office • Full-Time • 2+ years exp • San Francisco
C#
C#
.NET
AI/ML
Axolotl
DeepSpeed
Knowledge Distillation
LLM
Multimodal AI
PyTorch
Transformers
GPT-5
Hugging Face
Post-training
SFT
DevOps
GitHub
Analytics
ETL/ELT
Apply
$120k – $180k per year • In office • Full-Time • 5+ years exp • San Francisco
Go
TypeScript
C#
JavaScript
C#
.NET
gRPC for .NET
AI/ML
DeepSeek
Llama
LLM
Frontend
Lighthouse
Next.js
React.js
Recharts
Tailwind CSS
tRPC
visx
D3.js
DevOps
CI/CD
Docker
gRPC
Terraform
WebSockets
Design
Figma
QA
Chrome DevTools
Apply
$293k – $385k per year • In office • Full-Time • San Francisco
AI/ML
OpenAI
Apply
$180k – $260k per year • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • San Francisco
Python
AI/ML
ChatGPT
OpenAI
OpenAI Codex
Cybersecurity
FedRAMP
Apply
$180k – $210k per year • Equity • In office • Full-Time • San Francisco
Node JS
JavaScript
Databases
PostgreSQL
DevOps
PagerDuty
Web3
TRM Labs
Management
Slack
Apply
$252k – $335k per year • Remote/Hybrid • Full-Time • 8+ years exp • San Francisco
AI/ML
ChatGPT
Human-in-the-Loop
OpenAI
OpenAI Codex
DevOps
SLI/SLO/SLA
Apply
$223k – $424k per year (Estimated) • In office • Bachelor's Degree • San Francisco
AI/ML
AI Agents
LLM
Recommender Systems
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.