369,078open jobs
9,456companies
47,986added this week
Browse all
Salary
$161k – $354k per year (Estimated)
Location
In office (London)
Seniority
Staff · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Perplexity AI is an American company that builds an answer engine combining live web search with large language models to return sourced, conversational responses instead of a list of links. Its products span a consumer assistant on web and mobile, the Comet browser, enterprise search over internal documents and the Sonar developer API that exposes the same grounded retrieval stack. Founded in 2022 in San Francisco by former researchers and engineers from OpenAI, Meta and Databricks, the company is backed by NVIDIA, IVP, New Enterprise Associates and SoftBank.

We are looking for an AI Inference Engineer to join our growing team. We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL.

Responsibilities:

  • New models support. Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway.

  • GPU kernels migration to CuTe DSL. Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow.

  • Rust-native serving runtime. Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic.

  • Performance optimisation. Profile and fix bottlenecks from network ingress through continuous batching and GPU kernels interleaving.

  • Reliability and observability. Build dashboards, alerts, and automated remediation so we catch regressions before users do. Respond to and learn from production incidents.

Who we're looking for:

  • Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar). Any other deep systems programming experience is a plus.

  • You understand modern LLM architectures and are able to bring them up reliably in a production environment.

  • You've built and operated production distributed systems under real load - ideally performance-critical ones.

  • Comfortable working across languages and layers: Rust for the serving runtime, Python for model code, CUDA/CuteDSL for kernels.

  • You own problems end-to-end. You can read a research paper on Monday, write a kernel on Wednesday, and debug a production incident on Friday.

  • Self-directed. You do well in fast-moving environments where the path forward isn't laid out for you.

Nice-to-have:

  • ML compilers and framework internals: PyTorch internals, torch.compile, custom operators.

  • Distributed GPU communication: NCCL, NVLink, InfiniBand, RDMA libraries, model/tensor parallelism.

  • Low-precision inference: INT8/FP8/FP4 quantization, mixed-precision serving.

  • Profiling and debugging tools: Nsight Compute/Systems, CUDA-GDB, PTX/SASS analysis.

  • Container orchestration: Kubernetes, GPU scheduling, autoscaling inference workloads.

Qualifications:

  • 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems.

  • Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow).

  • Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores).

  • Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation).

Final offer amounts are determined by multiple factors including experience and expertise.

Equity: In addition to the base salary, equity may be part of the total compensation package.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
369,078 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
London
$23k – $61k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Bengaluru
C#
Python
C#
.NET
Databases
Oracle
Apply
$25k – $69k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Bengaluru • Pune
Java
Python
AI/ML
AI Agents
AWS Bedrock
Fine-tuning
Google ADK
LangGraph
LLM
LoRA
RAG
Semantic Search
LangChain
PEFT
A2A
Amazon SageMaker
AWS Strands Agents
NIST AI RMF
Semantic Search
Model Context Protocol
DevOps
Amazon EC2
Amazon EKS
AWS
Azure
CI/CD
CloudFormation
Docker
GCP
Git
GitOps
gRPC
Kubernetes
OpenTelemetry
Rest API
Terraform
Vector
Amazon S3
IAM
GitLab
Apply
$34k – $82k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru • Pune
Java
Python
Python
Asyncio
FastAPI
Databases
Amazon Aurora
AI/ML
AWS Bedrock
Google ADK
LangChain
LangGraph
LLM
RAG
A2A
LLM Guardrails
AI Agents
Model Context Protocol
DevOps
Amazon EKS
AWS
Envoy
Kubernetes
API Gateway
IAM
Cybersecurity
Zero Trust
Apply
$23k – $62k per year (Estimated) • In office • Full-Time • PhD • Pune
C++
Java
Python
SQL
Java
Spring Boot
Spring Cloud
Spring Security
Python
Django
SQLAlchemy
Databases
Apache Kafka
Redis
DevOps
CI/CD
GitLab CI
GitLab
Apply
AI Engineer 1 day ago
$26k – $108k per year (Estimated) • In office • Full-Time • Pune
Python
COBOL
Python
FastAPI
COBOL
IBM MQ
Databases
DynamoDB
AI/ML
AutoGen
AWS Bedrock
Claude
CrewAI
Embeddings
LangChain
LangGraph
Prompt Engineering
RAG
Edge AI
LLM Guardrails
AI Agents
DevOps
Amazon EKS
AWS
AWS Lambda
CI/CD
Docker
Kubernetes
Amazon CloudWatch
Amazon ECS
Amazon S3
API Gateway
IAM
Apply
$275k – $375k per year • In office • Full-Time • 15+ years exp • San Francisco
AI/ML
Perplexity
Web3
Rollup
Apply
$300k – $405k per year • In office • Full-Time • 8+ years exp • San Francisco
Go
Python
Rust
AI/ML
Perplexity
DevOps
Platform Engineering
Apply
$180k – $300k per year • In office • Full-Time • 3+ years exp • San Francisco • New York
TypeScript
JavaScript
AI/ML
Perplexity
Frontend
GSAP
Tailwind CSS
Analytics
A/B Testing
Design
Figma
Framer
Management
Slack
Apply
$200k – $250k per year • In office • Full-Time • San Francisco
AI/ML
LLM
Perplexity
RAG
Apply
$200k – $400k per year • Remote/Hybrid • Full-Time • 4+ years exp • San Francisco • New York
Kotlin
Rust
TypeScript
AI/ML
Perplexity
AI Agents
Apply
$65k – $155k per year (Estimated) • In office • Internship • Bachelor's Degree • London
Go
JavaScript
Ruby
Scala
Apply
In office • Internship • 1+ year exp • Bachelor's Degree • London
Go
JavaScript
Ruby
Scala
Apply
$221k – $370k per year • In office • Full-Time • 5+ years exp • London
AI/ML
OpenAI
Apply
$72k – $136k per year (Estimated) • Equity • Remote/Hybrid • 5+ years exp • London
Python
SQL
Databases
Snowflake
AI/ML
AI Agents
DevOps
Kibana
Marketing
Salesforce
Apply
$77k – $148k per year (Estimated) • Equity • In office • Master's Degree • London
JavaScript
Python
Scala
AI/ML
AI Agents
DevOps
GitHub
Apply
See all jobs
This is one of many
369,078 more open roles from verified company boards, updated every day.