368,941open jobs
9,452companies
47,951added this week
Browse all
Salary
$150k – $300k per year
Location
In office (San Francisco)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Prime Intellect is an artificial intelligence infrastructure company headquartered in San Francisco, California, and founded in 2023. The company provides a decentralized platform for training, evaluating, and deploying large-scale AI models, featuring tools for reinforcement learning, agent development, and a global compute marketplace. It operates globally by aggregating computing resources from various providers to enable researchers and developers to build open-source models and autonomous agents.

Own Your Intelligence

Prime Intellect is building the open superintelligence stack: the infrastructure frontier AI labs build internally, made available to every ambitious AI team.

Our platform, Lab, unifies compute, environments, evaluations, secure sandboxes, high-performance training, and deployment into one full-stack system for post-training at frontier scale - from SFT and RL to tool use, agent workflows, and continuously improving production models. We are building open frontier AI: open-source models trained end to end for long-horizon tasks like autonomous research, and the full-stack platform our own research team uses to build them. The next generation of AI companies, enterprises, and research teams do not just need more GPUs. They need the ability to turn their own workflows, tools, data, and feedback loops into superintelligence they own.

Prime Intellect has raised $150M in total funding from Founders Fund, Radical Ventures, NVIDIA, and exceptional AI, infrastructure, and enterprise operators - including Andrej Karpathy, Dwarkesh Patel, and leaders and founders from Ramp, Perplexity, Harvey, Mercor, Zapier, Datadog, Cognition, OpenAI, Thinking Machines, Together AI, SemiAnalysis, LangChain, Browserbase, Cloudflare, Sierra, Databricks, Airbnb, OpenRouter, Standard Intelligence, Fleet, Core Auto, and more. We are looking for people who want to build at the intersection of frontier research, real infrastructure, and go-to-market for a category that does not fully exist yet.

Role Impact

You'll help build our hosted training platform - the product that lets users launch LoRA and full fine-tuning runs on managed GPU clusters with a single API call or a few clicks. The role spans the developer-facing platform and the underlying Kubernetes-based training infrastructure that runs the jobs.

Core Technical Responsibilities

Hosted Training Infrastructure

  • Design and operate Kubernetes-based training and inference orchestration across multi-cluster, multi-cloud GPU fleets

  • Build and maintain Helm charts that compose trainers, inference servers, environment servers, and supporting services into reproducible "Training stacks"

  • Develop the Python control-plane agents that watch pods, report run state to the platform, and keep clusters in sync

  • Implement scheduling and autoscaling for heterogeneous hardware (H100/H200/B200) using KEDA, LeaderWorkerSet, taints/tolerations, and gang scheduling

  • Run a tight GitOps workflow - every change ships through PRs, Helm values, and CI

  • Build node-local model caches, checkpoint pipelines, and shared storage for fast cold starts

  • Operate the observability stack (Prometheus, Grafana, Loki, DCGM) and make GPU cluster debugging fast

Platform Development

  • Build the developer-facing surfaces for hosted training: job submission, live run monitoring, logs, metrics, model/adapter management, comparisons

  • Develop FastAPI backend services and REST APIs that bridge the platform to running clusters

  • Build real-time monitoring and debugging tools (streaming logs, step-level metrics, failure analysis)

  • Ship product UI in Next.js / React / TypeScript with shadcn, Tailwind, tRPC, and TanStack Query

Research Bridge

  • Interface with the RL trainer, inference servers, and environment servers running inside our clusters

  • Productize new training capabilities (new model architectures, RL algorithms, modes)

Technical Requirements

We're looking for engineers who are fluent across three areas - you don't need to be the world's best at any one, but you should have real depth in all three and a clear point of view on how they connect.

AI & GPU Landscape

  • Strong working knowledge of the modern AI stack - open model families, finetuning techniques (LoRA, QLoRA, full FT, RLHF/RLAIF), inference engines (vLLM, SGLang, TensorRT-LLM)

  • Familiarity with GPU hardware tradeoffs (H100 / H200 / B200, NVLink, interconnects, memory hierarchy) and what they mean for training and inference workloads

  • Understanding of distributed training fundamentals (data/tensor/pipeline/expert parallelism, NCCL, multi-node scheduling)

  • Awareness of what's happening at the frontier - new models, training methods, infra patterns - and the ability to translate that into product decisions

Kubernetes & Infrastructure

  • Strong Kubernetes operations experience - Helm, CRDs, operators, KEDA, gang scheduling, GPU operator

  • Comfortable debugging real production clusters (kubectl, pod lifecycle, node issues, networking)

  • Cloud platform experience (GCP preferred - GCS, GKE, Cloud Run, Cloud Tasks)

  • Infrastructure automation (Helm, Terraform, Ansible) and a GitOps mindset

  • Observability: Prometheus, Grafana, Loki, OpenTelemetry, DCGM

  • Linux fundamentals: networking, namespaces, performance tuning

Programming & Platform

  • Strong Python backend development (FastAPI, async, SQLAlchemy)

  • Comfortable building Python control-plane agents that talk to Kubernetes APIs

  • Modern frontend development (TypeScript, React/Next.js, Tailwind, shadcn) - enough to ship product surfaces end-to-end

  • REST and tRPC API design

  • Experience building developer tools, dashboards, and live-monitoring UIs

What We Offer

  • Cash compensation $150K-$300K with significant equity

  • Flexible work arrangement (remote or San Francisco office)

  • Full visa sponsorship and relocation support

  • Professional development budget for courses and conferences

  • Regular team off-sites and conference attendance

  • Opportunity to shape the future of decentralized AI development

Growth Opportunity

You'll join a team of experienced engineers and researchers working on cutting-edge problems in AI infrastructure. We believe in open development and encourage team members to contribute to the broader AI community through research and open-source work.

We value potential over perfection - if you're passionate about democratizing AI development and have experience in either platform or infrastructure development (ideally both), we want to talk to you.

Ready to help shape the future of AI? Apply now and join us in our mission to make powerful AI models accessible to everyone.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,941 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$22k – $38k per year (net) • In office • Full-Time • 3+ years exp • Astana
Python
TypeScript
JavaScript
Python
Alembic
Celery
FastAPI
Litestar
Pydantic
SQLAlchemy
structlog
Databases
ClickHouse
ElasticSearch
MinIO
NATS
Neo4j
PostgreSQL
Redis
AI/ML
LLM
OpenAI
Pydantic AI
Frontend
Mantine
React Hook Form
React.js
Recharts
shadcn/ui
Tailwind CSS
TanStack Router
TanStack Table
Vite
Zustand
Radix UI
DevOps
Amazon S3
Ansible
Docker
Git
GitLab
GitLab CI
Grafana
gRPC
Terraform
QA
Playwright
Pytest
Apply
$96k – $218k per year (Estimated) • Equity • In office • Full-Time • 8+ years exp • Toronto
Python
Databases
Databricks
Snowflake
AI/ML
AI Agents
AWS Bedrock
AWS Bedrock AgentCore
LLM
LLM Evaluation
DevOps
AWS
CI/CD
GCP
Apply
$143k – $173k per year • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Wiesbaden
Python
DevOps
Ansible
Terraform
Cybersecurity
Defense in Depth
Wireshark
Zero Trust
Apply
$128k – $173k per year • In office • Full-Time • 7+ years exp • United States
Python
SQL
Databases
Databricks
Snowflake
DevOps
AWS
SLI/SLO/SLA
Analytics
Power BI
Tableau
Management
Confluence
Jira
Apply
$62k – $142k per year (Estimated) • In office • Full-Time • 6+ years exp • PhD • Madrid
Apex
JavaScript
Python
TypeScript
Databases
Databricks
Google BigQuery
Snowflake
AI/ML
Agentforce
AI Agents
Claude
Cursor
LangChain
LlamaIndex
LLM
Prompt Engineering
Marketing
Salesforce
Apply
$180k – $350k per year • Equity • In office • Full-Time • 5+ years exp • San Francisco
Python
Rust
Databases
Databricks
AI/ML
LangChain
OpenRouter
Perplexity
Together AI
OpenAI
Post-training
Red Teaming
SFT
AI Agents
Function Calling
DevOps
Cloudflare
Datadog
eBPF
GCP
Kubernetes
Service Mesh
Cybersecurity
Falco
FedRAMP
ISO 27001
SOC 2
Threat Modeling
Zero Trust
Management
Zapier
Apply
$150k – $300k per year • Equity • In office • Full-Time • San Francisco
TypeScript
JavaScript
Databases
Databricks
AI/ML
AI Agents
DSPy
LangChain
LangGraph
LLM
OpenRouter
Perplexity
Ray
Reinforcement Learning
RLHF
SGLang
Synthetic Data
Together AI
vLLM
GRPO
LLM Evaluation
OpenAI
Post-training
SFT
Function Calling
Model Context Protocol
Frontend
Next.js
React.js
DevOps
Cloudflare
Datadog
Docker
Grafana
Kubernetes
Prometheus
Terraform
Management
Zapier
Apply
$184k – $372k per year (Estimated) • In office • Full-Time • San Francisco
Python
Rust
TypeScript
JavaScript
Python
FastAPI
Databases
Databricks
AI/ML
LangChain
OpenRouter
Perplexity
Together AI
OpenAI
Post-training
SFT
TPU
Function Calling
Frontend
Next.js
React.js
Tailwind CSS
DevOps
Ansible
Cloudflare
Datadog
GCP
Grafana
Kubernetes
Prometheus
Rest API
Terraform
WebSockets
Management
Zapier
Apply
$150k – $300k per year • Equity • In office • Full-Time • San Francisco
Go
Python
Rust
TypeScript
JavaScript
Python
FastAPI
Databases
Databricks
AI/ML
LangChain
OpenRouter
Perplexity
Together AI
OpenAI
Post-training
SFT
TPU
Function Calling
Frontend
Next.js
React.js
Tailwind CSS
DevOps
Ansible
Cloudflare
Datadog
GCP
Grafana
Kubernetes
Prometheus
Rest API
Terraform
WebSockets
Management
Zapier
Apply
$151k – $274k per year (Estimated) • Remote • Bachelor's Degree • San Francisco
Databases
Databricks
AI/ML
LangChain
LLM
OpenRouter
Perplexity
Together AI
OpenAI
Post-training
SFT
Function Calling
DevOps
Cloudflare
Datadog
Management
Zapier
Apply
$130k – $500k per year • Equity • In office • Full-Time • 5+ years exp • San Francisco
AI/ML
AI Agents
Claude
Claude Code
Copilot
Cursor
Function Calling
Human-in-the-Loop
LLM Guardrails
DevOps
GitHub
Apply
$89k – $193k per year (Estimated) • In office • Full-Time • 3+ years exp • High School Diploma • San Francisco
Apply
Founding Engineer 3 hours ago
$120k – $150k per year • In office • Full-Time • 3+ years exp • San Francisco
C++
Go
Rust
Chips/EDA
KiCad
Apply
$300k – $475k per year • Remote/Hybrid • Full-Time • 5+ years exp • San Francisco
Apply
$350k – $475k per year • In office • Full-Time • 4+ years exp • San Francisco • New York
C++
Python
C++
PyTorch C++
AI/ML
PyTorch
Ray
Reinforcement Learning
RLHF
DPO
InfiniBand
NCCL
Post-training
PPO
TPU
DevOps
Kubernetes
SLURM
SRE
Apply
See all jobs
This is one of many
368,941 more open roles from verified company boards, updated every day.