368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$106k – $252k per year (Estimated)
Location
In office (Cambridge)
Seniority
Principal
Employment
Full-Time
Overview
Company
Impact
Profile match
Lila Sciences is an artificial intelligence enterprise developing an autonomous platform designed to accelerate discovery across the life, chemical, and materials sciences. Headquartered in Cambridge, Massachusetts, the company pairs generative AI models with automated robotic laboratories to formulate hypotheses, execute physical experiments, and analyze results. Its integrated system aims to streamline the scientific method, facilitating the rapid creation of novel therapeutics, advanced materials, and industrial compounds for global commercial partners.

Your Impact at LILA

The Staff/Principal DevOps Engineer - AI Inference will drive the design, implementation, and optimization of infrastructure purpose-built for serving machine learning models at scale. This role bridges platform engineering, site reliability, and ML infrastructure, building the systems that power low-latency, high-throughput inference across GPU clusters and cloud accelerators. You will collaborate with ML engineers, research scientists, and software engineers to build inference platforms that serve models reliably to production users while maximizing compute efficiency.

What You'll Be Building

  • GPU/accelerator infrastructure on Kubernetes: scheduling, resource isolation, multi-tenant GPU sharing, device plugins, and topology-aware placement for inference workloads
  • Model serving platforms using frameworks such as vLLM, Triton Inference Server, TGI, or custom serving stacks with optimized batching, caching, and request routing
  • Intelligent request routing and load balancing across heterogeneous accelerator fleets (NVIDIA GPUs, AWS Inferentia/Trainium) to maximize utilization and minimize latency
  • Autoscaling systems that dynamically match inference compute supply with demand across production, research, and experimental workloads
  • Production-grade deployment pipelines for ML models: canary rollouts, A/B testing, model versioning, and safe rollback across multi-region deployments
  • Infrastructure-as-code with Terraform and Helm for GPU-accelerated EKS clusters, including node pools, spot/on-demand strategies, and accelerator-specific networking
  • Observability and performance optimization: GPU utilization monitoring, inference latency profiling, token throughput dashboards, and SLO/SLI tracking for model endpoints
  • CI/CD pipelines for model artifacts: container image builds with CUDA/driver dependencies, model registry integration, and automated inference benchmarking in CI
  • AWS cloud infrastructure for ML: EKS with GPU node groups, EC2 accelerated instances (P4/P5, Inf2, Trn1), S3 model storage, EFA/high-bandwidth networking, and IAM least privilege
  • Cost optimization and capacity planning: right-sizing accelerator instances, spot instance strategies for inference, and fleet-wide efficiency reporting

What You'll Need to Succeed

  • Expertise in DevOps, SRE, or Platform Engineering with significant experience operating GPU/accelerator infrastructure at scale
  • Deep experience with Kubernetes for ML workloads: GPU scheduling, resource quotas, node affinity, and accelerator device management
  • Strong proficiency deploying to AWS using infrastructure-as-code (Terraform, Helm) with hands-on experience managing GPU-based compute (EKS, EC2 P-series/Inf/Trn instances)
  • Experience with model serving infrastructure: inference servers, request batching, KV-cache optimization, or LLM serving frameworks
  • Strong understanding of networking for distributed inference: high-bandwidth interconnects, NCCL, VPC/PrivateLink, and load balancing at L4/L7
  • Strong proficiency in Python for automation, tooling, and integration with ML frameworks

Bonus Points For

  • Experience with LLM inference optimization: continuous batching, speculative decoding, quantization (GPTQ, AWQ, FP8), tensor parallelism, and pipeline parallelism
  • Hands-on experience with multiple accelerator families (NVIDIA A100/H100, AWS Inferentia2, Trainium, AMD MI300X) and maintaining hardware-agnostic serving infrastructure
  • Multi-region deployment experience with geographic routing and failover for latency-sensitive inference endpoints
  • Proficiency in Rust or Go for performance-critical infrastructure components
  • SRE practices for ML systems: chaos engineering on GPU workloads, incident management, capacity modeling for bursty inference traffic
  • Experience with model registries, artifact versioning, and ML supply chain security
  • Observability platform expertise: building custom metrics for token-level throughput, time-to-first-token, and per-request GPU memory profiling
  • Prior startup/high-growth experience balancing velocity with reliability in rapidly scaling AI systems

Compensation

We offer competitive base compensation with bonus potential and generous early-stage equity. Your final offer will reflect your background, expertise, and expected impact.

U.S. Benefits. Full-time U.S. employees receive a comprehensive benefits program including medical, dental, and vision coverage; employer-paid life and disability insurance; flexible time off with generous company wide holidays; paid parental leave; an educational assistance program; commuter benefits, including bike share memberships for office based employees; and a company subsidized lunch program.

International Benefits. Full-time employees outside the U.S. receive a comprehensive benefits program tailored to their region. USD salary ranges apply only to U.S.-based positions; international salaries are set to local market.

Expected Base Salary Range

$192,000—$272,000 USD

About LILA

Lila Sciences is building Scientific Superintelligence™ to solve humankind's greatest challenges. We believe science is the most inspiring frontier for AI. Rather than hard-coding expert knowledge into tools, LILA builds systems that can learn for themselves.

LILA combines advanced AI models with proprietary AI Science Factory™ instruments into an operating system for science that executes the entire scientific method autonomously, accelerating discovery at unprecedented speed, scale, and impact across medicine, materials, and energy. Learn more at www.lila.ai.

Guided by our core values of truth, trust, curiosity, grit, and velocity, we move with startup speed while tackling problems of historic importance. If this sounds like an environment you'd love to work in, even if you don't meet every qualification listed above, we encourage you to apply.

We’re All In

Lila Sciences is committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status.

Information you provide during your application process will be handled in accordance with our Candidate Privacy Policy.

A Note to Agencies

Lila Sciences does not accept unsolicited resumes from any source other than candidates. The submission of unsolicited resumes by recruitment or staffing agencies to Lila Sciences or its employees is strictly prohibited unless contacted directly by Lila Science’s internal Talent Acquisition team. Any resume submitted by an agency in the absence of a signed agreement will automatically become the property of Lila Sciences, and Lila Sciences will not owe any referral or other fees with respect thereto.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Cambridge
$105k – $252k per year • Remote • Full-Time • 18+ years exp • Bachelor's Degree
Python
Java
Java
Gradle
DevOps
Ansible
AWS
CI/CD
CloudFormation
Configuration Management
Docker
GitHub Actions
GitLab CI
Helm
Jenkins
Kubernetes
Platform Engineering
Terraform
GitHub
GitLab
Cybersecurity
Sonatype Nexus IQ
Management
Confluence
Jira
Apply
$54k – $175k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$35k – $113k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$35k – $116k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$133k – $161k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Westminster
C++
Python
DevOps
CI/CD
SpaceTech
NASA cFS
Apply
$146k – $305k per year (Estimated) • In office • Full-Time • Cambridge
TypeScript
JavaScript
AI/ML
AI Agents
Frontend
React.js
Analytics
ETL/ELT
Apply
$146k – $305k per year (Estimated) • In office • Full-Time • Cambridge
TypeScript
JavaScript
AI/ML
NumPy
Pandas
SciPy
AI Agents
Frontend
React.js
Analytics
ETL/ELT
Apply
$43k – $136k per year (Estimated) • In office • Full-Time • 3+ years exp • Cambridge
Cybersecurity
CAPA
Apply
$58k – $126k per year (Estimated) • In office • Full-Time • Cambridge
C++
Python
DevOps
Git
Rest API
Cybersecurity
CAPA
Design
AutoCAD
SolidWorks
IoT
MQTT
OPC UA
Apply
$154k – $330k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Cambridge
Python
SQL
TypeScript
Python
FastAPI
AI/ML
AI Agents
Copilot
LLM
RAG
Context Engineering
Knowledge Graph
Model Context Protocol
DevOps
AWS
CI/CD
CloudFormation
GitHub Actions
Kubernetes
Terraform
GitHub
Apply
$113k – $189k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • Cambridge
Python
AI/ML
LLM
NLP
NumPy
Pandas
Prompt Engineering
PyTorch
Hugging Face
OpenAI
DevOps
Git
Apply
$113k – $189k per year • In office • Full-Time • 8+ years exp • Master's Degree • Cambridge
Python
AI/ML
LLM
NLP
NumPy
Pandas
Prompt Engineering
PyTorch
Hugging Face
OpenAI
DevOps
Git
Apply
$113k – $189k per year • In office • Full-Time • 15+ years exp • Bachelor's Degree • Cambridge
Python
AI/ML
LLM
NLP
NumPy
Pandas
Prompt Engineering
PyTorch
Hugging Face
OpenAI
DevOps
Git
Apply
$82k – $206k per year • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Cambridge
C++
MATLAB
Python
MATLAB
Simulink
Apply
$75k – $245k per year • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Cambridge
C++
MATLAB
Python
MATLAB
Simulink
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.