369,078open jobs
9,456companies
47,986added this week
Browse all
Salary
$170k – $245k per year
Location
In office (San Francisco, Palo Alto)
Employment
Full-Time
Overview
Company
Impact
Profile match
Anyscale is an artificial intelligence infrastructure company headquartered in San Francisco, California, and founded in 2019 by the creators of Ray at the UC Berkeley RISELab. The company offers a managed platform for running Ray, the open source framework used to scale model training, batch inference, and reinforcement learning across clusters. It sells to machine learning engineering teams that need to move distributed AI workloads from laptops to production without rebuilding their stack.

About Anyscale

At Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray, a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise, and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world.

With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.

Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.

About the role

As a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale. This is an incredibly critical role to Anyscale as it allows us to achieve a market leading position for AI infrastructure.

As part of this role, you will

  • Iterate very quickly with product teams to ship the end to end solutions for Batch and Online inference at high scale which will be used by open-source Ray users and customers of Anyscale

  • Work across the stack integrating Ray Data and LLM engine providing optimizations achieving low cost solutions for large scale ML inference

  • Integrate with Open source software like vLLM, work closely with the community to adopt these techniques in Anyscale solutions, and also contribute improvements to open source

  • Follow the latest state-of-the-art in the open source and the research community, implementing and extending best practices

We'd love to hear from you if you have

  • Familiarity with running ML inference at large scale with high throughput and low latency

  • Familiarity with deep learning and deep learning frameworks (e.g. PyTorch)

  • Solid understanding of distributed systems, ML inference challenges

Bonus points!

  • ML Systems knowledge

  • Experience using Ray

  • Work closely with community on LLM engines like vLLM, TensorRT-LLM

  • Contributions to deep learning frameworks (PyTorch, TensorFlow)

  • Contributions to deep learning compilers (Triton, TVM, MLIR)

  • Prior experience working on GPUs / CUDA

Compensation

At Anyscale, we take a market-based approach to compensation. We are data-driven, transparent, and consistent. As the market data changes over time, the target salary for this role may be adjusted.

This role is also eligible to participate in Anyscale's Equity and Benefits offerings, including the following:

  • Stock Options

  • Healthcare plans, with premiums covered by Anyscale at 99% for both employees and dependents

  • 401k Retirement Plan

  • Education & Wellbeing Stipend

  • Paid Parental Leave

  • Fertility Benefits

  • Paid Time Off

  • Commute reimbursement

  • 100% of in-office meals covered

Anyscale Inc. is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law.

Anyscale Inc. is an E-Verify company and you may review theNotice of E-Verify Participationand the Right to Work posters in English and Spanish

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
369,078 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$98k – $164k per year • In office • Full-Time • 5+ years exp • Master's Degree • United States
Python
AI/ML
LLM
Multimodal AI
PyTorch
TensorFlow
Knowledge Graph
OpenAI
Time Series Forecasting
DevOps
GitHub
Apply
$41k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
C#
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
AI/ML
Copilot
LLM
OpenAI
Frontend
Angular
GraphQL
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Rest API
Apply
$80k – $175k per year • In office • Full-Time • Toronto
Python
AI/ML
AWS Bedrock
Claude
Copilot
LLM
Prompt Engineering
RAG
Context Engineering
AI Agents
DevOps
AWS
CI/CD
Splunk
GitHub
Apply
Sr. UX Designer 8 hours ago
$76k – $161k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Toronto
Python
AI/ML
AI Agents
Hallucination
Human-in-the-Loop
LLM
LLM Guardrails
Design
Figma
Apply
$223k – $424k per year (Estimated) • In office • Bachelor's Degree • San Francisco
AI/ML
AI Agents
LLM
Recommender Systems
Apply
$226k – $283k per year • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco
AI/ML
Fine-tuning
LLM
MLFlow
PyTorch
Ray
OpenAI
Post-training
Model Context Protocol
DevOps
OpenTelemetry
Apply
$180k – $210k per year • In office • Full-Time • 8+ years exp • San Francisco
AI/ML
Ray
OpenAI
Cybersecurity
CVSS
SBOM
Threat Modeling
Apply
$200k – $240k per year • Remote/Hybrid • Full-Time • 8+ years exp • San Francisco
AI/ML
Ray
OpenAI
DevOps
AWS
Azure
Kubernetes
Cybersecurity
ISO 27001
SOC 2
Apply
IT Specialist 5 days ago
$171k – $211k per year • Equity • In office • Full-Time • San Francisco
AI/ML
Ray
CoreWeave
OpenAI
DevOps
Azure
GCP
GitHub
Cybersecurity
Okta
Management
Google Workspace
Apply
$200k – $240k per year • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • San Francisco
Go
Python
AI/ML
Ray
OpenAI
TPU
DevOps
AWS
Azure
GCP
Grafana
Kubernetes
Prometheus
Apply
$130k – $500k per year • Equity • In office • Full-Time • 5+ years exp • San Francisco
AI/ML
AI Agents
Claude
Claude Code
Copilot
Cursor
Function Calling
Human-in-the-Loop
LLM Guardrails
DevOps
GitHub
Apply
$89k – $193k per year (Estimated) • In office • Full-Time • 3+ years exp • High School Diploma • San Francisco
Apply
Founding Engineer 4 hours ago
$120k – $150k per year • In office • Full-Time • 3+ years exp • San Francisco
C++
Go
Rust
Chips/EDA
KiCad
Apply
$300k – $475k per year • Remote/Hybrid • Full-Time • 5+ years exp • San Francisco
Apply
$350k – $475k per year • In office • Full-Time • 4+ years exp • San Francisco • New York
C++
Python
C++
PyTorch C++
AI/ML
PyTorch
Ray
Reinforcement Learning
RLHF
DPO
InfiniBand
NCCL
Post-training
PPO
TPU
DevOps
Kubernetes
SLURM
SRE
Apply
See all jobs
This is one of many
369,078 more open roles from verified company boards, updated every day.