368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$220k – $485k per year
Location
In office (San Francisco, New York, Palo Alto)
Seniority
Staff · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Perplexity AI is an American company that builds an answer engine combining live web search with large language models to return sourced, conversational responses instead of a list of links. Its products span a consumer assistant on web and mobile, the Comet browser, enterprise search over internal documents and the Sonar developer API that exposes the same grounded retrieval stack. Founded in 2022 in San Francisco by former researchers and engineers from OpenAI, Meta and Databricks, the company is backed by NVIDIA, IVP, New Enterprise Associates and SoftBank.

We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL - and we need another engineer to join us.

What you will work on

Examples of real work the team does:

  • New models support. Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway.

  • GPU kernels migration to CuTe DSL. Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow.

  • Rust-native serving runtime. Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic.

  • Performance optimisation. Profile and fix bottlenecks from network ingress through continuous batching and GPU kernel interleaving.

  • Reliability and observability. Build dashboards, alerts, and automated remediation so we catch regressions before users do. Respond to and learn from production incidents.

Who we're looking for

  • Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar). Any other deep systems programming experience is a plus.

  • You understand modern LLM architectures and are able to bring them up reliably in a production environment.

  • You've built and operated production distributed systems under real load - ideally performance-critical ones.

  • Comfortable working across languages and layers: Rust for the serving runtime, Python for model code, CUDA/CuteDSL for kernels.

  • You own problems end-to-end. You can read a research paper on Monday, write a kernel on Wednesday, and debug a production incident on Friday.

  • Self-directed. You do well in fast-moving environments where the path forward isn't laid out for you.

Good if you touched any of

  • ML compilers and framework internals: PyTorch internals, torch.compile, custom operators.

  • Distributed GPU communication: NCCL, NVLink, InfiniBand, RDMA libraries, model/tensor parallelism.

  • Low-precision inference: INT8/FP8/FP4 quantization, mixed-precision serving.

  • Profiling and debugging tools: Nsight Compute/Systems, CUDA-GDB, PTX/SASS analysis.

  • Container orchestration: Kubernetes, GPU scheduling, autoscaling inference workloads.

Qualifications

  • 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems.

  • Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow).

  • Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores).

  • Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation).

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$143k – $258k per year (Estimated) • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
JavaScript
Python
TypeScript
Python
pySpark
AI/ML
Prompt Engineering
Spark
DevOps
AWS
Azure
CI/CD
GCP
Git
Jenkins
GitHub
GitLab
Analytics
ETL/ELT
Apply
$163k – $434k per year • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
JavaScript
Python
TypeScript
Python
pySpark
AI/ML
Prompt Engineering
Spark
DevOps
AWS
Azure
CI/CD
GCP
Git
Jenkins
GitHub
GitLab
Analytics
ETL/ELT
Apply
$32k – $83k per year (Estimated) • In office • Full-Time • 5+ years exp • Gurgaon
Python
Python
pySpark
Databases
Microsoft Fabric
AI/ML
Spark
DevOps
Azure
Apply
Remote • Full-Time • 3+ years exp • Prague
Bash
PowerShell
Python
DevOps
Splunk
Cybersecurity
Crowdstrike
IBM QRadar
Apply
$44k – $95k per year (Estimated) • In office • Full-Time • Tokyo • Fukuoka
Java
PHP
Python
TypeScript
JavaScript
Frontend
Next.js
React.js
Management
Power Apps
Power Automate
Apply
$275k – $375k per year • In office • Full-Time • 15+ years exp • San Francisco
AI/ML
Perplexity
Web3
Rollup
Apply
$300k – $405k per year • In office • Full-Time • 8+ years exp • San Francisco
Go
Python
Rust
AI/ML
Perplexity
DevOps
Platform Engineering
Apply
$180k – $300k per year • In office • Full-Time • 3+ years exp • San Francisco • New York
TypeScript
JavaScript
AI/ML
Perplexity
Frontend
GSAP
Tailwind CSS
Analytics
A/B Testing
Design
Figma
Framer
Management
Slack
Apply
$200k – $250k per year • In office • Full-Time • San Francisco
AI/ML
LLM
Perplexity
RAG
Apply
$200k – $400k per year • Remote/Hybrid • Full-Time • 4+ years exp • San Francisco • New York
Kotlin
Rust
TypeScript
AI/ML
Perplexity
AI Agents
Apply
$180k – $210k per year • Equity • In office • Full-Time • San Francisco
Node JS
JavaScript
Databases
PostgreSQL
DevOps
PagerDuty
Web3
TRM Labs
Management
Slack
Apply
$252k – $335k per year • Remote/Hybrid • Full-Time • 8+ years exp • San Francisco
AI/ML
ChatGPT
Human-in-the-Loop
OpenAI
OpenAI Codex
DevOps
SLI/SLO/SLA
Apply
$223k – $424k per year (Estimated) • In office • Bachelor's Degree • San Francisco
AI/ML
AI Agents
LLM
Recommender Systems
Apply
$160k – $283k per year • Equity • In office • 5+ years exp • San Francisco
AI/ML
AI Agents
Apply
$185k – $385k per year • Remote/Hybrid • Full-Time • 5+ years exp • San Francisco
JavaScript
Python
Databases
MySQL
PostgreSQL
AI/ML
OpenAI
Frontend
React.js
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.