838,726open jobs
53,672companies
140,652added this week
Browse all
Salary
$150k – $220k per year
Location
Remote (United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 27, 2026. First seen by Alion on Sep 25, 2026. RunPod scores B on the Alion truth index.

Overview
Company
Impact
Profile match
RunPod is a cloud computing company headquartered in Mount Laurel, New Jersey, and founded in 2022. The company provides a specialized AI developer cloud featuring on-demand GPU instances, serverless GPU endpoints, and multi-node GPU clusters for training and inference. It operates a global infrastructure across more than 30 regions, serving over one million developers with scalable compute resources for machine learning and generative AI workloads.

Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at an inflection point for AI infrastructure, and we're building the platform the next generation of developers will depend on.

We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale.

Learn more in our CEO's funding announcement: https://www.runpod.io/blog/one-million-developers.

We're looking for a ML Systems Engineer, Inference. We want Runpod to be the best place in the world to run LLM inference, meaning the fastest and the most cost-efficient. You'll lead that effort. You'll own LLM serving performance end to end. That means measuring it, understanding it, and improving it across models, hardware generations, and workloads. The work you ship will show up directly in the latency and cost our customers experience. This is a hands-on engineering role for someone who likes finding the real bottleneck and fixing it, then turning that fix into something that runs reliably in production.

Responsibilities

  • Define how we measure inference performance, including throughput, time to first token, inter-token latency, and cost per token, and build the tooling that makes those measurements rigorous and repeatable.

  • Profile and diagnose performance problems across the serving stack, from scheduling and memory management down to kernels and interconnect.

  • Improve serving efficiency for large, state-of-the-art models on single-node and multi-node GPU deployments.

  • Turn what you learn into production-ready runtimes, configurations, and defaults that customers benefit from automatically.

  • Work closely with product and infrastructure teams to shape how inference is offered on Runpod.

  • Keep up with the fast-moving inference ecosystem, including the open-source community, and decide what's worth adopting, what's worth building, and what's worth contributing back.

  • Trace bottlenecks in the serving engine/runtime and implement fixes when configuration tuning is not enough.

Requirements

  • 5+ years of professional system engineering experience.

  • Deep, hands-on experience with vLLM, SGLang (or a comparable serving engine) in production or at serious benchmark scale.

  • Strong software engineering skills in Python. You're comfortable working in large, performance-critical codebases.

  • A solid understanding of what drives LLM inference performance: batching, memory, parallelism, and the trade-offs between latency and throughput.

  • Experience with modern inference optimization techniques such as quantization, speculative decoding, or distributed serving.

  • Rigor in benchmarking and performance analysis, plus comfort with GPU profiling tools.

  • The ability to explain your results clearly in writing and turn them into decisions.

Preferred

  • Experience writing or tuning GPU kernels in CUDA or Triton.

  • Contributions to inference or ML systems projects.

  • Experience with multi-node GPU systems and high-speed networking.

  • Experience at a company where inference cost and latency were core business metrics.

What You’ll Receive:

  • The competitive base pay for this position ranges from ($150,000 - $220,000). This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate’s experience, qualifications, and location

  • Meaningful equity in a fast-growing company- everyone on the team receives stock options - your impact drives our growth, and you share in the upside.

  • Generous medical, dental & vision plans

  • Flexible PTO- take the time you need to recharge

  • Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication

  • Join a passionate team on the cutting edge of AI infrastructure - where culture, learning, and ownership are at the heart of how we scale.

  • $1,200 Home Office & Equipment Stipend- We set you up for success from day one with gear and support to create your ideal workspace

Runpod is committed to maintaining a workplace free from discrimination and upholding the principles of equality and respect for all individuals. We believe that diversity in all its forms enhances our team. As an equal opportunity employer, Runpod is committed to creating an inclusive workforce at every level. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, marital status, protected veteran status, disability status, or any other characteristic protected by law. We welcome every qualified candidate eligible to work in the United States; however, we are currently unable to sponsor employment visas.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
838,726 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
In your city
$53k – $68k per year • Remote (Ireland) • 6+ years exp • PhD • Dublin
Python
C++
Julia
MATLAB
Apply
$74k – $219k per year • In office • Full-Time • 12+ years exp • Associate's Degree • New York • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
Claude
Model Context Protocol
AI Agents
Anthropic
Context Engineering
Knowledge Graph
Edge AI
Tool Use
Apply
$99k – $136k per year • Remote (United States) • Full-Time • 3+ years exp • PhD • Evanston
AI/ML
NLP
LLM
Apply
$90k – $123k per year • In office • Full-Time • 2+ years exp • PhD • Skokie
Python
SQL
DevOps
Git
Apply
≈ $111k – $243k per year (Estimated) • Remote (Mexico) • 5+ years exp • Bachelor's Degree
Python
JavaScript
AI/ML
Copilot
Claude
ChatGPT
Prompt Engineering
AI Agents
RAG
Copilot Studio
Management
ServiceNow
Power Automate
ITSM
Apply
≈ $50k – $112k per year (Estimated) • In office • Contractor • Montreal
Python
SQL
AI/ML
Copilot
Analytics
Power BI
Microsoft Excel
Management
Power Automate
SharePoint
Apply
≈ $33k – $77k per year (Estimated) • In office • Poland
Python
Bash
DevOps
RTOS
CI/CD
Git
Linux
Apply
≈ $82k – $162k per year (Estimated) • Equity • In office • Full-Time • 7+ years exp • Bachelor's Degree • Charlotte • Atlanta
Python
JavaScript
SQL
Management
Agile
Marketing
Salesforce
Marketo
Apply
Atlantbh Internship 2 days ago
≈ $15k – $29k per year (Estimated) • Remote (likely Bosnia and Herzegovina) • Internship • Sarajevo
Python
JavaScript
Java
TypeScript
Node JS
Bash
Java
Maven
Spring Boot
Node JS
Axios
Databases
PostgreSQL
AI/ML
NLP
Tokenization
Sentiment Analysis
Interpretability
Machine Learning
Frontend
Angular
React.js
Mobile
JUnit
DevOps
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Nginx
Linux
Design
Figma
Management
Trello
Jira
Agile
Scrum
QA
Playwright
Postman
Pytest
Apply
AI Ops Engineer 2 days ago
≈ $29k – $74k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Chennai
Python
Databases
Google BigQuery
BigQuery
AI/ML
LLM
Hallucination
LLMOps
Machine Learning
DevOps
Terraform
GCP
CI/CD
GitOps
Docker
Kubernetes
Platform Engineering
Self-Healing
Google GKE
Google Cloud Run
AIOps
IAM
Apply
$180k – $240k per year • Equity • Remote (United States) • Full-Time
Python
Go
Rust
Databases
MinIO
AI/ML
Fine-tuning
RunPod
DevOps
Datadog
Prometheus
Kubernetes
Grafana
Amazon S3
Linux
Management
Slack
Apply
$160k – $280k per year • Equity • Remote (United States) • Full-Time • 6+ years exp
AI/ML
RunPod
Marketing
HubSpot
Apply
$120k – $160k per year • Equity • Remote (United States) • Full-Time • 3+ years exp
Python
Go
JavaScript
Node JS
Bash
Node JS
Commander.js
AI/ML
AI Agents
RunPod
InfiniBand
DevOps
Datadog
Prometheus
Docker
Grafana
SLI/SLO/SLA
HPC
Linux
Management
Slack
Apply
$100k – $160k per year • Equity • Remote (APAC, United States) • Full-Time • 3+ years exp • Bachelor's Degree
Python
Go
JavaScript
SQL
Node JS
Python
Flask
Django
Databases
MySQL
PostgreSQL
AI/ML
Fine-tuning
AI Agents
PyTorch
LLM
RunPod
Machine Learning
Frontend
React.js
DevOps
Docker
Ubuntu
Linux
TCP/IP
DNS
Apply
$130k – $300k per year • Equity • Remote (APAC, United States) • Full-Time • 8+ years exp
AI/ML
RunPod
Apply
See all jobs
This is one of many
838,726 more open roles from verified company boards, updated every day.