368,657open jobs
9,442companies
50,883added this week
Browse all
Salary
$150k – $300k per year
Location
Remote/Hybrid (San Francisco, United States)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Prime Intellect is an artificial intelligence infrastructure company headquartered in San Francisco, California, and founded in 2023. The company provides a decentralized platform for training, evaluating, and deploying large-scale AI models, featuring tools for reinforcement learning, agent development, and a global compute marketplace. It operates globally by aggregating computing resources from various providers to enable researchers and developers to build open-source models and autonomous agents.

Own Your Intelligence

Prime Intellect is building the open superintelligence stack: the infrastructure frontier AI labs build internally, made available to every ambitious AI team.

Our platform, Lab, unifies compute, environments, evaluations, secure sandboxes, high-performance training, and deployment into one full-stack system for post-training at frontier scale - from SFT and RL to tool use, agent workflows, and continuously improving production models. We are building open frontier AI: open-source models trained end to end for long-horizon tasks like autonomous research, and the full-stack platform our own research team uses to build them. The next generation of AI companies, enterprises, and research teams do not just need more GPUs. They need the ability to turn their own workflows, tools, data, and feedback loops into superintelligence they own.

Prime Intellect has raised $150M in total funding from Founders Fund, Radical Ventures, NVIDIA, and exceptional AI, infrastructure, and enterprise operators - including Andrej Karpathy, Dwarkesh Patel, and leaders and founders from Ramp, Perplexity, Harvey, Mercor, Zapier, Datadog, Cognition, OpenAI, Thinking Machines, Together AI, SemiAnalysis, LangChain, Browserbase, Cloudflare, Sierra, Databricks, Airbnb, OpenRouter, Standard Intelligence, Fleet, Core Auto, and more. We are looking for people who want to build at the intersection of frontier research, real infrastructure, and go-to-market for a category that does not fully exist yet.

Role Impact

This is a hybrid position spanning cloud LLM serving, LLM inference optimization and RL systems. You will be working on advancing our ability to evaluate and serve models trained with our RL Lab at scale. The two key areas are:

  • Building the infrastructure to serve LLMs efficiently at scale.

  • Optimization and integration of inference systems into our RL training stack.

Core Technical Responsibilities

LLM Serving

  • Multi-tenant LLM Serving: Build a multi-tenant LLM serving platform that operates across our cloud GPU fleets.

  • GPU-Aware Scheduling: Design placement and scheduling algorithms for heterogeneous accelerators.

  • Resilience & Failover: Implement multi-region/zone failover and traffic shifting for resilience and cost control.

  • Autoscaling & Routing: Build autoscaling, routing, and load balancing to meet throughput/latency SLOs.

  • Model Distribution: Optimize model distribution and cold-start times across clusters.

Inference Optimization & Performance

  • Framework Development: Integrate and contribute to LLM inference frameworks such as vLLM, SGLang, TensorRT-LLM.

  • Parallelism and Configuration Tuning: Optimize configurations for tensor/pipeline/expert parallelism, prefix caching, memory management and other axes for maximum performance.

  • End-to-End Performance: Profile kernels, memory bandwidth and transport; apply techniques such as quantization and speculative decoding.

  • Perf Suites: Develop reproducible performance suites (latency, throughput, context length, batch size, precision).

  • RL Integration: Embed and optimize distributed inference within our RL stack.

Platform & Tooling

  • CI/CD: Establish CI/CD with artifact promotion, performance gates, and reproducible builds.

  • Observability: Build metrics, logs, tracing; structured incident response and SLO management.

  • Docs & Collaboration: Document architectures, playbooks, and API contracts; mentor and collaborate cross-functionally.

Technical Requirements

Required Experience

  • Building ML Systems at Scale: 3+ years building and running large-scale ML/LLM services with clear latency/availability SLOs.

  • Inference Backends: Hands-on with at least one of vLLM, SGLang, TensorRT-LLM.

  • Distributed Serving Infra: Familiarity with distributed and disaggregated serving infrastructure such as NVIDIA Dynamo.

  • Inference Internals: Deep understanding of prefill vs. decode, KV-cache behavior, batching, sampling, speculative decoding, parallelism strategies.

  • Full-Stack Debugging: Comfortable debugging CUDA/NCCL, drivers/kernels, containers, service mesh/networking, and storage, owning incidents end-to-end.

Infrastructure Skills

  • Python: Systems tooling and backend services.

  • PyTorch: LLM Inference engine development and integration, deployment readiness.

  • Cloud & Automation: AWS/GCP service experience, cloud deployment patterns.

  • Kubernetes: Running infrastructure at scale with containers on Kubernetes.

  • GPU & Networking: Architecture, CUDA runtime, NCCL, InfiniBand; GPU-aware bin-packing and scheduling across heterogeneous fleets.

Nice to Have

  • Kernel-Level Optimization: Familiarity with CUDA/Triton kernel development; Nsight Systems/Compute profiling.

  • Systems Performance Languages: Rust, C++.

  • Data & Observability: Kafka/PubSub, Redis, gRPC/Protobuf; Prometheus/Grafana, OpenTelemetry; reliability patterns.

  • Infra & Config Automation: Terraform/Ansible, infrastructure-as-code, reproducible environments

  • Open Source: Contributions to serving, inference, or RL infrastructure projects.

What We Offer

  • Cash Compensation Range of $150-300k with significant equity incentives

  • Flexible work arrangement (remote or San Francisco office)

  • Full visa sponsorship and relocation support

  • Professional development budget

  • Regular team off-sites and conference attendance

  • Opportunity to shape decentralized AI and RL at Prime Intellect

Growth Opportunity

You'll join a team of experienced engineers and researchers working on cutting-edge problems in AI infrastructure. We believe in open development and encourage team members to contribute to the broader AI community through research and open-source contributions.

We value potential over perfection. If you're passionate about democratizing AI development, we want to talk to you.

Ready to help shape the future of AI? Apply now and join us in our mission to make powerful AI models accessible to everyone.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,657 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$25k – $42k per year • Equity 0–0.2% • Remote • Full-Time • 3+ years exp
Bash
Go
JavaScript
Python
TypeScript
DevOps
AWS
Azure
CI/CD
Datadog
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Incident Management
Kubernetes
Platform Engineering
Prometheus
Terraform
Amazon CloudWatch
GitHub
GitLab
IAM
Cybersecurity
Least Privilege
Apply
$100k – $210k per year • Equity 0–0.5% • Remote • Full-Time • 3+ years exp • San Francisco
Bash
Go
JavaScript
Python
TypeScript
DevOps
AWS
Azure
CI/CD
Datadog
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Incident Management
Kubernetes
Platform Engineering
Prometheus
Terraform
Amazon CloudWatch
GitHub
GitLab
IAM
Cybersecurity
Least Privilege
Apply
$49k – $173k per year (Estimated) • In office • 2+ years exp • Bachelor's Degree • Kfar Saba
Java
Kotlin
Python
SQL
C#
C#
.NET
DevOps
AWS
CI/CD
Apply
$100k – $200k per year • Equity 0.5–5% • In office • Full-Time • 1+ year exp • New York
Python
TypeScript
JavaScript
Python
FastAPI
Databases
DynamoDB
PostgreSQL
AI/ML
Claude
LLM
OpenAI
AI Agents
Frontend
Next.js
Tailwind CSS
React.js
DevOps
AWS
Docker
Vercel
GitHub
Management
Slack
Apply
$14k – $31k per year (Estimated) • Remote/Hybrid • 3+ years exp • Moscow
JavaScript
DevOps
CI/CD
Git
GitLab CI
GitLab
Management
Confluence
Jira
QA
Playwright
Postman
Apply
$180k – $350k per year • Equity • In office • Full-Time • 5+ years exp • San Francisco
Python
Rust
Databases
Databricks
AI/ML
LangChain
OpenRouter
Perplexity
Together AI
OpenAI
Post-training
Red Teaming
SFT
AI Agents
Function Calling
DevOps
Cloudflare
Datadog
eBPF
GCP
Kubernetes
Service Mesh
Cybersecurity
Falco
FedRAMP
ISO 27001
SOC 2
Threat Modeling
Zero Trust
Management
Zapier
Apply
$150k – $300k per year • Equity • In office • Full-Time • San Francisco
TypeScript
JavaScript
Databases
Databricks
AI/ML
AI Agents
DSPy
LangChain
LangGraph
LLM
OpenRouter
Perplexity
Ray
Reinforcement Learning
RLHF
SGLang
Synthetic Data
Together AI
vLLM
GRPO
LLM Evaluation
OpenAI
Post-training
SFT
Function Calling
Model Context Protocol
Frontend
Next.js
React.js
DevOps
Cloudflare
Datadog
Docker
Grafana
Kubernetes
Prometheus
Terraform
Management
Zapier
Apply
$184k – $372k per year (Estimated) • In office • Full-Time • San Francisco
Python
Rust
TypeScript
JavaScript
Python
FastAPI
Databases
Databricks
AI/ML
LangChain
OpenRouter
Perplexity
Together AI
OpenAI
Post-training
SFT
TPU
Function Calling
Frontend
Next.js
React.js
Tailwind CSS
DevOps
Ansible
Cloudflare
Datadog
GCP
Grafana
Kubernetes
Prometheus
Rest API
Terraform
WebSockets
Management
Zapier
Apply
$150k – $300k per year • Equity • In office • Full-Time • San Francisco
Go
Python
Rust
TypeScript
JavaScript
Python
FastAPI
Databases
Databricks
AI/ML
LangChain
OpenRouter
Perplexity
Together AI
OpenAI
Post-training
SFT
TPU
Function Calling
Frontend
Next.js
React.js
Tailwind CSS
DevOps
Ansible
Cloudflare
Datadog
GCP
Grafana
Kubernetes
Prometheus
Rest API
Terraform
WebSockets
Management
Zapier
Apply
$150k – $300k per year • In office • Full-Time • San Francisco
Python
TypeScript
JavaScript
Python
FastAPI
SQLAlchemy
Databases
Databricks
AI/ML
Fine-tuning
LangChain
LLM
LoRA
OpenRouter
Perplexity
QLoRA
RLHF
SGLang
TensorRT
TensorRT-LLM
Together AI
vLLM
PEFT
NCCL
NVLink
OpenAI
Post-training
SFT
Function Calling
Frontend
Next.js
React.js
shadcn/ui
Tailwind CSS
tRPC
Radix UI
DevOps
Ansible
Cloudflare
Datadog
GCP
GitOps
Google Cloud Run
Google GKE
Grafana
Helm
KEDA
kubectl
Kubernetes
Loki
OpenTelemetry
Prometheus
Rest API
Terraform
Management
Zapier
Apply
$170k – $220k per year • Equity 1–2.8% • In office • Full-Time • 3+ years exp • San Francisco
Python
SQL
Python
Django
AI/ML
AI Agents
Context Engineering
LLM
LLM Evaluation
RAG
Apply
$173k – $314k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • San Francisco
Apex
JavaScript
Node JS
Python
SQL
TypeScript
Apex
Lightning Web Components
AI/ML
Agentforce
AI Agents
Claude
Claude Code
Copilot
Cursor
LLM
RAG
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
Grafana
gRPC
Kubernetes
New Relic
Prometheus
Splunk
Marketing
Salesforce
QA
Cypress
JMeter
k6
Locust
Playwright
Postman
Rest-Assured
Selenium
Apply
Senior ML Engineer 2 hours ago
$149k – $224k per year • In office • Full-Time • 5+ years exp • Master's Degree • San Francisco • Washington • Palo Alto
Python
Python
pySpark
Databases
Apache Kafka
AI/ML
AI Agents
Agentforce
Airflow
Anomaly Detection
Feature Store
Flink
Ray
Red Teaming
Spark
DevOps
CI/CD
Docker
Kubernetes
Cybersecurity
MITRE ATT&CK
Marketing
Salesforce
Apply
In office • Internship • 1+ year exp • Bachelor's Degree • San Francisco
Go
JavaScript
Ruby
Scala
Apply
$360k – $530k per year • In office • Full-Time • Bachelor's Degree • San Francisco
MATLAB
Python
MATLAB
Simulink
AI/ML
OpenAI
Robotics
Digital Twin
Apply
See all jobs
This is one of many
368,657 more open roles from verified company boards, updated every day.