698,241open jobs
41,057companies
105,127added this week
Browse all
Salary
$250k – $395k per year
Location
Remote (San Francisco, United States)
Employment
Full-Time
Overview
Company
Impact
Profile match
URun is infrastructure for real-time AI video, built for interactive creation where teams can generate, steer, and evolve scenes with no delays or compromise.

The problem we saw

Most AI infrastructure is built for batch: send a query, wait, get a response, reset. Powerful, but transactional. AI is becoming interactive - sessions that hold state, models that stay alive between turns, generation that responds as it runs - and the infrastructure to deliver that at scale doesn't really exist yet.

The bottleneck isn't the models anymore. It's the infrastructure underneath them.

What we're building to fix it

uRun is the inference cloud for interactive AI: the compute layer that makes real-time, stateful inference possible at scale. We came out of stealth in April 2026, are backed by top-tier investors, and are founded by Keegan McCallum, who scaled inference infrastructure for some of the most demanding generative AI workloads in production.

We're an infrastructure company. We build the layer that model labs, builders, and research teams ship on top of.

Where you come in

Performance is uRun's core differentiator. We're not chasing incremental gains - we're building infrastructure that runs 10-100x faster than the status quo. As our ML Performance Engineer, you will be the person who makes that true.

This is a founding technical hire. You will write custom CUDA kernels, push GPU utilization to its limits, and own inference latency end-to-end across the stack. You will work directly with the founding team on the hardest performance problems in production AI infrastructure - and your fingerprints will be on everything we ship.

What you'll actually be doing day-to-day

  • Write custom CUDA kernels that unlock performance headroom unavailable through off-the-shelf frameworks

  • Optimize model inference end-to-end, targeting sub-50ms latency across our inference platform

  • Drive 10x performance improvements across the stack: memory bandwidth, kernel fusion, operator scheduling, and beyond

  • Implement zero-copy distributed memory optimizations across multi-GPU and multi-node environments

  • Own GPU utilization and memory management, squeezing every available FLOP out of the hardware we run

  • Profile, benchmark, and instrument the full inference pipeline to find and eliminate bottlenecks systematically

  • Set the performance engineering bar for the team: define what fast looks like and build the tooling to measure it

What skills you need for the journey

  • Deep, hands-on CUDA expertise: you have written custom kernels in production, not just called into cuBLAS

  • Strong background in model inference and post-training optimization at scale

  • Fluency in GPU memory hierarchy, warp scheduling, kernel fusion, and hardware-aware algorithm design

  • Experience profiling and benchmarking complex inference pipelines: you know where the time goes and how to get it back

  • Able to operate at the frontier with minimal guidance - you identify the problem, design the approach, and ship the fix

Things that will give you an edge

  • Public work in GPU optimization or inference efficiency - open source contributions, a published paper, or a side project that shows your depth (vLLM, Flash-Attention, TensorRT-LLM, PyTorch, or equivalent)

  • Experience with hardware-aware optimization frameworks: CuTe, Triton, TileLang, or similar

  • Familiarity with distributed memory and communication primitives: NCCL, InfiniBand, NVLink, RoCE

  • Contributions to or deep familiarity with PyTorch Distributed, Ray core, or similar systems

  • Experience optimizing for video generation or other high-throughput, latency-sensitive generative workloads

  • Prior work at an inference-focused company or research lab pushing the boundary of what GPU hardware can do

What you'll get in return

Competitive salary and meaningful equity in an early-stage AI infrastructure company. The band above is our target; for an exceptional candidate we'll go higher. Equity is real - you're early, and the grant reflects that.

  • Health, dental, and vision

  • 401(k) - company-supported retirement savings

  • FSA/HSA - flexible spending accounts for healthcare costs

  • Paid time off - we trust you to manage your time

  • Top-tier tooling - access to the best AI tools available: Claude, Codex, Kimi, and whatever else helps you move faster

  • MacBook Pro and AirPods - the hardware you need, on us

How we work (and what that feels like day-to-day)

We build the stage, not the show. We're an infrastructure company, a developer-tools company, and a production partner for model labs, and focus is a deliberate choice we've made and hold to.

Day-to-day, that means a small team, a high bar, and real ownership. You won't wait for permission or inherit a backlog of someone else's decisions, in a founding security role, the function is what you make it.

It also means ambiguity: priorities shift, not everything is documented, and you'll often be the person who decides what "secure enough, for now" means. That suits some people and not others, and we'd rather you know that before you apply.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
698,241 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
Data Scientist 7 hours ago
$18k – $43k per year (Estimated) • In office • Full-Time • 3+ years exp • Master's Degree
Python
JavaScript
TypeScript
SQL
Databases
PostgreSQL
pgvector
AI/ML
LangChain
Claude
OpenCV
Claude Code
LlamaIndex
YOLO
Scikit-learn
Prompt Engineering
Computer Vision
NLP
Detectron2
TensorFlow
NumPy
PyTorch
RAG
BERT
Torchvision
Hallucination
Sentiment Analysis
OpenAI
Hugging Face
Machine Learning
Frontend
Tailwind CSS
React.js
Framer Motion
DevOps
Docker
Analytics
A/B Testing
Management
n8n
Zapier
Apply
$138k – $297k per year (Estimated) • In office • Full-Time • 7+ years exp • PhD • San Francisco • Chicago • New York
AI/ML
Claude
AI Agents
OpenAI Codex
Agentforce
Management
Slack
Marketing
Salesforce
Apply
Senior Engineer, QE 3 days ago
In office • 5+ years exp • Pune
JavaScript
TypeScript
AI/ML
Copilot
Cursor
Claude
Agentic Workflows
Frontend
Webpack
React.js
Lighthouse
DevOps
GCP
GitHub Actions
CI/CD
Jenkins
AWS
Docker
Kubernetes
Shift-Left
Buildkite
GitLab
Cybersecurity
Shift-Left Security
QA
Selenium
Cypress
Playwright
Jest
Vitest
k6
Apply
$26k – $68k per year (Estimated) • In office • Contractor
Python
AI/ML
Claude
Model Context Protocol
AI Agents
LLM
Anthropic
Multi-Agent Systems
Management
n8n
Apply
$125k – $260k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree
Python
Java
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
LangGraph
LangChain
TensorFlow
PyTorch
Time Series Forecasting
Machine Learning
Apply
$200k – $400k per year • Remote • Full-Time • San Francisco
AI/ML
Claude
Apply
$200k – $350k per year • Remote • Full-Time • 7+ years exp • San Francisco
Python
Go
JavaScript
TypeScript
Node JS
AI/ML
Claude
DevOps
WebRTC
WebSockets
Kubernetes
Platform Engineering
Apply
$250k – $350k per year • Remote • Full-Time • 7+ years exp • San Francisco
Python
Go
JavaScript
TypeScript
Node JS
AI/ML
Claude
DevOps
Terraform
GCP
OpenTelemetry
WebRTC
Prometheus
Pulumi
Azure
AWS
Kubernetes
Grafana
SLI/SLO/SLA
IAM
Apply
$200k – $350k per year • Remote • Full-Time • San Francisco
AI/ML
Claude
vLLM
SGLang
TensorRT
TensorRT-LLM
Triton
Feature Store
NCCL
InfiniBand
DevOps
GCP
SLURM
Azure
AWS
Kubernetes
SLI/SLO/SLA
Apply
$59k – $126k per year (Estimated) • Remote/Hybrid • Contractor • 1+ year exp • Bachelor's Degree • San Francisco
Management
Microsoft Office
Apply
$97k – $188k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • New York • Boston • San Francisco • Washington • Seattle
Management
Outlook
Microsoft Office
Apply
$133k – $338k per year • Remote • Full-Time • 12+ years exp • Associate's Degree • New York • Milwaukee • Dallas • Columbus • Kirkland
Python
Databases
Neo4j
Amazon Neptune
AI/ML
Airflow
Prompt Engineering
Multimodal AI
AI Agents
NLP
TensorFlow
PyTorch
LLM
Knowledge Graph
Machine Learning
DevOps
GCP
Azure
AWS
Analytics
ETL/ELT
Apache NiFi
Apply
$123k – $317k per year • In office • Full-Time • 10+ years exp • Master's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
Management
Agile
Marketing
Salesforce
Apply
$99k – $205k per year (Estimated) • In office • Full-Time • 3+ years exp • High School Diploma • San Francisco
DevOps
SLI/SLO/SLA
Windows
Management
ITIL
Service Desk
Apply
See all jobs
This is one of many
698,241 more open roles from verified company boards, updated every day.