368,657open jobs
9,442companies
50,883added this week
Browse all
Salary
$126k – $299k per year (Estimated)
Location
Remote/Hybrid (London, United Kingdom)
Seniority
Staff · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Perplexity AI is an American company that builds an answer engine combining live web search with large language models to return sourced, conversational responses instead of a list of links. Its products span a consumer assistant on web and mobile, the Comet browser, enterprise search over internal documents and the Sonar developer API that exposes the same grounded retrieval stack. Founded in 2022 in San Francisco by former researchers and engineers from OpenAI, Meta and Databricks, the company is backed by NVIDIA, IVP, New Enterprise Associates and SoftBank.

We are looking for an AI Infra engineer to join our growing team. We work with Kubernetes, Slurm, Python, C++, PyTorch, and primarily on AWS. As an AI Infrastructure Engineer, you will be partnering closely with our Inference and Research teams to build, deploy, and optimize our large-scale AI training and inference clusters.

Responsibilities

  • Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads

  • Manage and optimize Slurm-based HPC environments for distributed training of large language models

  • Develop robust APIs and orchestration systems for both training pipelines and inference services

  • Implement resource scheduling and job management systems across heterogeneous compute environments

  • Benchmark system performance, diagnose bottlenecks, and implement improvements across both training and inference infrastructure

  • Build monitoring, alerting, and observability solutions tailored to ML workloads running on Kubernetes and Slurm

  • Respond swiftly to system outages and collaborate across teams to maintain high uptime for critical training runs and inference services

  • Optimize cluster utilization and implement autoscaling strategies for dynamic workload demands

Qualifications

  • Strong expertise in Kubernetes administration, including custom resource definitions, operators, and cluster management

  • Hands-on experience with Slurm workload management, including job scheduling, resource allocation, and cluster optimization

  • Experience with deploying and managing distributed training systems at scale

  • Deep understanding of container orchestration and distributed systems architecture

  • High level familiarity with LLM architecture and training processes (Multi-Head Attention, Multi/Grouped-Query, distributed training strategies)

  • Experience managing GPU clusters and optimizing compute resource utilization

Required Skills

  • Expert-level Kubernetes administration and YAML configuration management

  • Proficiency with Slurm job scheduling, resource management, and cluster configuration

  • Python and C++ programming with focus on systems and infrastructure automation

  • Hands-on experience with ML frameworks such as PyTorch in distributed training contexts

  • Strong understanding of networking, storage, and compute resource management for ML workloads

  • Experience developing APIs and managing distributed systems for both batch and real-time workloads

  • Solid debugging and monitoring skills with expertise in observability tools for containerized environments

Preferred Skills

  • Experience with Kubernetes operators and custom controllers for ML workloads

  • Advanced Slurm administration including multi-cluster federation and advanced scheduling policies

  • Familiarity with GPU cluster management and CUDA optimization

  • Experience with other ML frameworks like TensorFlow or distributed training libraries

  • Background in HPC environments, parallel computing, and high-performance networking

  • Knowledge of infrastructure as code (Terraform, Ansible) and GitOps practices

  • Experience with container registries, image optimization, and multi-stage builds for ML workloads

Required Experience

  • Demonstrated experience managing large-scale Kubernetes deployments in production environments

  • Proven track record with Slurm cluster administration and HPC workload management

  • Previous roles in SRE, DevOps, or Platform Engineering with focus on ML infrastructure

  • Experience supporting both long-running training jobs and high-availability inference services

  • Ideally, 3-5 years of relevant experience in ML systems deployment with specific focus on cluster orchestration and resource management

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,657 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
London
$169k – $253k per year • Equity • In office • Full-Time • 8+ years exp • Bachelor's Degree • Boulder
JavaScript
TypeScript
AI/ML
AI Agents
Frontend
GraphQL
React.js
DevOps
AWS
Azure
GCP
Google GKE
Helm
Kubernetes
Apply
$58k – $136k per year (Estimated) • Remote/Hybrid • Full-Time • Guadalajara
Bash
Node JS
Python
JavaScript
Python
FastAPI
Databases
RabbitMQ
Frontend
Next.js
React.js
DevOps
ArgoCD
AWS
CI/CD
Datadog
Docker
Git
GitHub
GitHub Actions
GitOps
Grafana
Helm
Jenkins
Karpenter
KEDA
Kubernetes
OpenTelemetry
Prometheus
Terraform
Apply
HLS specialist 2 hours ago
In office • Full-Time • Israel
DevOps
AWS
Azure
Apply
$66k – $123k per year • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Corvallis
Python
SQL
Databases
Snowflake
AI/ML
AI Agents
dbt
DevOps
AWS
GitHub
Analytics
Power BI
Tableau
Apply
$60k – $108k per year • Remote/Hybrid • Full-Time • Bachelor's Degree • United States
PowerShell
Python
DevOps
AWS
Azure
IAM
Splunk
Cybersecurity
Crowdstrike
ISO 27001
Microsoft Defender
Microsoft Entra ID
Microsoft Sentinel
NIST CSF
Qualys Cloud Platform
Apply
$275k – $375k per year • In office • Full-Time • 15+ years exp • San Francisco
AI/ML
Perplexity
Web3
Rollup
Apply
$300k – $405k per year • In office • Full-Time • 8+ years exp • San Francisco
Go
Python
Rust
AI/ML
Perplexity
DevOps
Platform Engineering
Apply
$180k – $300k per year • In office • Full-Time • 3+ years exp • San Francisco • New York
TypeScript
JavaScript
AI/ML
Perplexity
Frontend
GSAP
Tailwind CSS
Analytics
A/B Testing
Design
Figma
Framer
Management
Slack
Apply
$200k – $250k per year • In office • Full-Time • San Francisco
AI/ML
LLM
Perplexity
RAG
Apply
$200k – $400k per year • Remote/Hybrid • Full-Time • 4+ years exp • San Francisco • New York
Kotlin
Rust
TypeScript
AI/ML
Perplexity
AI Agents
Apply
$65k – $155k per year (Estimated) • In office • Internship • Bachelor's Degree • London
Go
JavaScript
Ruby
Scala
Apply
In office • Internship • 1+ year exp • Bachelor's Degree • London
Go
JavaScript
Ruby
Scala
Apply
$221k – $370k per year • In office • Full-Time • 5+ years exp • London
AI/ML
OpenAI
Apply
$72k – $136k per year (Estimated) • Equity • Remote/Hybrid • 5+ years exp • London
Python
SQL
Databases
Snowflake
AI/ML
AI Agents
DevOps
Kibana
Marketing
Salesforce
Apply
$77k – $148k per year (Estimated) • Equity • In office • Master's Degree • London
JavaScript
Python
Scala
AI/ML
AI Agents
DevOps
GitHub
Apply
See all jobs
This is one of many
368,657 more open roles from verified company boards, updated every day.