708,188open jobs
42,107companies
99,812added this week
Browse all
Salary
$200k – $400k per year
Location
In office (Palo Alto)
Seniority
Middle · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match

About Us

We are a well-funded, early-stage AI lab focused on building the next generation of frontier multimodal AI models. Founded by former DeepMind researchers, including Andrew Dai, who was previously a leader on Gemini. Our team currently consists of 20 world-class scientists and engineers. We recently raised $55M in seed funding from Striker Ventures, Menlo Ventures, Altimeter Capital, and NVIDIA. We are tackling some of the hardest problems in artificial intelligence, and we are growing fast.

About the Role

We're looking for an infrastructure engineer to design, optimize, and scale the systems that serve our large multimodal models. Your work will make inference faster, more cost-effective, and more reliable, so our teams can focus on advancing model capabilities rather than managing bottlenecks.

Our focus is on performant, efficient inference, both to power real-world applications and to accelerate research. This role owns the infrastructure that ensures every deployment and evaluation runs smoothly at scale for our visual foundation models.

What You Will Do

  • Build low-latency, high-throughput inference serving systems for our large multimodal models

  • Design and implement techniques that improve latency, throughput, and efficiency, including quantization, batching, speculative decoding, and KV cache management

  • Optimize our codebase and GPU fleet to fully utilize hardware FLOPs, bandwidth, and memory

  • Implement multi-GPU and multi-node model parallelism for serving (tensor or pipeline parallel)

  • Build autoscaling and load balancing for production ML services

  • Establish standards for reliability, observability, and reproducibility across the inference stack

  • Collaborate with researchers to enable high-performance inference for novel architectures

Skills and Qualifications

Minimum qualifications:

  • 3+ years of experience building low-latency, high-throughput inference serving systems for large models

  • Strong knowledge of inference optimization techniques (quantization, batching, speculative decoding, KV cache management)

  • Hands-on experience with serving frameworks such as vLLM, TensorRT-LLM, Triton, or SGLang

  • Experience with multi-GPU/multi-node model parallelism for serving (tensor or pipeline parallel)

  • Strong systems programming skills; C++/CUDA a plus alongside Python

  • Experience with autoscaling and load balancing for production ML services

  • A track record of GPU cost optimization at scale

Preferred qualifications (strong candidates may have some, not all):

  • Experience serving multimodal (vision + language) models

  • Contributions to open-source ML or systems infrastructure projects (e.g., vLLM, SGLang, TensorRT-LLM, Triton)

  • A bias for action and comfort working across stacks and teams in an early-stage environment

Logistics

  • Location: This role is based on-site in Palo Alto, California.

  • Compensation: Depending on background, skills, and experience, the expected annual base salary range for this position is $200,000 - $400,000 USD, plus equity and benefits.

  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

  • Benefits: We offer health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Elorian AI is an equal opportunity employer. We are committed to building a diverse team and inclusive environment.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
708,188 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Palo Alto
$32k – $80k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Madrid
Python
TypeScript
AI/ML
Copilot
LangGraph
LangChain
Model Context Protocol
AI Agents
Langfuse
LLM
Machine Learning
DevOps
Rest API
CI/CD
Git
Docker
GitLab
QA
Playwright
Postman
Pytest
Apply
$39k – $86k per year (Estimated) • Remote • Full-Time • 8+ years exp • Bengaluru
Python
Go
AI/ML
Model Context Protocol
AI Agents
A2A
DevOps
Rest API
Splunk
Terraform
Ansible
OpenShift
GitHub Actions
Loki
Prometheus
CI/CD
GitOps
ArgoCD
Jenkins
AWS
Docker
Kubernetes
Grafana
OpenStack
Linux
Apply
$48k – $114k per year (Estimated) • In office • Full-Time • Bengaluru
Python
AI/ML
Cursor
LangGraph
LangChain
Claude Code
LlamaIndex
LoRA
Fine-tuning
Prompt Engineering
Function Calling
AI Agents
Arize Phoenix
Langfuse
PEFT
QLoRA
Transformers
LLM
RAG
LLMOps
GraphRAG
Human-in-the-Loop
Structured Outputs
Agentic Workflows
Tool Use
Model Distillation
Machine Learning
Chips/EDA
PoC Library
Apply
Remote/Hybrid • Full-Time • Hyderabad
Python
Databases
PostgreSQL
AI/ML
AI Agents
LLM
RAG
DevOps
GCP
GitHub Actions
CI/CD
Kubernetes
GitHub
IAM
Cybersecurity
Burp Suite
OWASP ZAP
SonarQube
Trivy
Semgrep
ISO 27001
CodeQL
OWASP Top 10
SOC 2
Least Privilege
SBOM
Dependabot
SIEM
OWASP
Apply
$84k – $182k per year (Estimated) • In office • Singapore
Python
JavaScript
AI/ML
AI Agents
DevOps
GitHub Actions
Azure
CI/CD
GitHub
IAM
Cybersecurity
Microsoft Sentinel
Least Privilege
Apply
General Interest 5 months ago
$123k – $262k per year (Estimated) • Remote/Hybrid • Full-Time • Palo Alto
AI/ML
Multimodal AI
Apply
Senior Director, Tax 12 hours ago
$273k – $376k per year • Remote/Hybrid • Full-Time • 15+ years exp • Bachelor's Degree • Palo Alto
Apply
$117k – $224k per year • In office • Full-Time • 4+ years exp • PhD • Palo Alto • Washington • San Francisco
Python
Java
TypeScript
Java
Maven
Databases
DynamoDB
AI/ML
Copilot
Cursor
Claude Code
AI Agents
LLM
OpenAI Codex
Agentforce
Human-in-the-Loop
LLM Guardrails
DevOps
CI/CD
Jenkins
AWS
Docker
Bazel
AWS Lambda
Amazon S3
Cybersecurity
Zero Trust
Apply
$240k – $300k per year • Remote/Hybrid • Full-Time • 6+ years exp • San Francisco • New York • Santa Monica • Los Angeles • Palo Alto
Cybersecurity
GDPR
Management
Microsoft Office
Apply
$45k – $108k per year (Estimated) • In office • Palo Alto
Python
SQL
Management
Jira
Apply
$166k – $248k per year • In office • 8+ years exp • PhD • Palo Alto
AI/ML
Fine-tuning
AI Agents
LLM Guardrails
DevOps
GCP
AWS
Marketing
X (Twitter)
LinkedIn
Instagram
Apply
See all jobs
This is one of many
708,188 more open roles from verified company boards, updated every day.