713,007open jobs
42,513companies
99,377added this week
Browse all
Salary
$240k – $290k per year
Location
Remote (United States)
Seniority
Staff · 5+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 23, 2026. First seen by Alion on Sep 22, 2026. Runway scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Runway is a mobile release management company headquartered in San Francisco, California, and founded in 2020. The company builds a platform that coordinates app store releases, automating build tracking, regression checks, staged rollouts, and the cross-team communication that surrounds a mobile launch. It integrates with the app stores, continuous integration systems, issue trackers, and chat tools that mobile teams already use, and went through Y Combinator.

We are building AI to simulate the world through merging art and science.

We believe that world models are at the frontier of progress in artificial intelligence. Language models alone won’t solve the world’s hardest problems - robotics, disease, scientific discovery. Real progress requires models that experience the world and learn from their mistakes, the same way that humans do. And this kind of trial and error can be massively accelerated when done in simulation, rather than in the real world.

World models offer the most clear path to general-purpose simulation, changing how stories are told, how scientific progress is made and how the next frontiers of humanity are reached.

Our team consists of creative, open minded, caring and ambitious people who are determined to change the world. We aspire to continuously build impossible things and our ability to do so relies on building an incredible team. If you are driven to do the same, we'd love to hear from you.

About the role

We're looking for an ML infrastructure engineer to own model evaluation at Runway, end to end. Every decision we make about a model - which checkpoint to keep training, what to ship to millions of users, which datamix and architecture shows the most promise - rests on evals. Today that work is spread across the organization. You'll turn it into one platform.

Your job will be to design and build the systems that generate samples at scale, score them with automated metrics and human annotations, track results across checkpoints and releases, and put the answers in front of researchers in minutes rather than days. You'll define what "better" means operationally - how we measure it, how confident we are, and how a result becomes a ship/no-ship decision.

You'll be embedded within research teams as a member of ML Platform, and you'll set the technical direction for evals across the company. This is an extremely high-leverage role: the quality of our models is bounded by how well we can measure them.

A peek at our technical stack

Our inference and training code is written in Python (PyTorch) and is cloud-native. We leverage multiple clouds and use a combination of Kubernetes native and home-grown tooling for efficient job orchestration.

We’re big users of agentic development and operations. We have a robust internal platform and provide broad access to the latest models and harnesses.

You’ll have access to best-in-class models, agents, GPUs, storage and cloud services you need to drive a world class evaluation pipeline that helps us deliver frontier models.

What you'll do

  • Own the evaluation platform end to end: the tooling and systems required to generate, annotate, review and adapt at frontier scale

  • Define the CLIs, APIs, GUIs and storage layers required to make this process seamless, fast, sophisticated and collaborative

  • Work directly with research teams on video, image, audio, agents, and robotics to understand what they need to measure and build it, then generalize the result into the platform

  • Help set the standards for how we evaluate models at Runway: reproducibility, metric definitions, reporting formats, and when a result is trustworthy enough to act on

  • Support broad adoption of the platform and its integration throughout Runway’s research efforts across training, production model serving and more

  • Contribute broadly as a member of the ML Platform team to tools and systems that help Runway train and serve frontier models

What you'll need

  • 5+ years of experience building ML infrastructure or data platforms in production environments, with at least some of that time spent on evaluation, experimentation, or benchmarking systems

  • Strong Python and PyTorch, and hands-on experience running large batch GPU workloads on Kubernetes

  • Experience designing data pipelines and storage for large volumes of media or model outputs, with attention to versioning and reproducibility

  • Familiarity with experimental statistics: paired comparisons, confidence intervals, multiple-comparison pitfalls, inter-rater agreement

  • Comfort building internal tools end to end, from the command line to the browser

  • Ability to lead a broad technical area: gather requirements from many teams, set direction, make tradeoffs, and drive a roadmap without waiting to be told

  • Familiarity with the full model development lifecycle: data, training, evaluation, serving

  • Self-starter who can work embedded with research teams and move fast

  • Strong systems thinking and pragmatic approach to production reliability

  • Humility and open mindedness; at Runway we love to learn from one another

Nice to have

  • Track record of building evaluation suites for generative models

  • Hands-on work with LLM- or VLM-as-judge pipelines

  • Experience with online experimentation platforms and connecting offline metrics to product outcomes

  • Prior work evaluating agents or robotics policies

Runway strives to recruit and retain exceptional talent from diverse backgrounds while ensuring pay equity for our team. Our salary ranges are based on competitive market rates for our size, stage and industry, and salary is just one part of the overall compensation package we provide.

There are many factors that go into salary determinations, including relevant experience, skill level and qualifications assessed during the interview process, and maintaining internal equity with peers on the team. The range shared below is a general expectation for the function as posted, but we are also open to considering candidates who may be more or less experienced than outlined in the job description. In this case, we will communicate any updates in the expected salary range.

Lastly, the provided range is the expected salary for candidates in the U.S. Outside of those regions, there may be a change in the range, which again, will be communicated to candidates.

Working at Runway

Great things come from great teams.We’d love to hear from you.

We’re committed to creating a space where our employees can bring their full selves to work and have equal opportunity to succeed. So regardless of race, gender identity or expression, sexual orientation, religion, origin, ability, age, veteran status, if joining this mission speaks to you, we encourage you to apply.

More about Runway

We're excited to be recognized as a best place to work:

Crain's | InHerSight | BuiltIn NYC | INC

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
713,007 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$127k – $190k per year • Equity • Remote/Hybrid • Full-Time • Bachelor's Degree • Seattle
Python
Go
Java
C++
Python
Flask
FastAPI
C++
TensorFlow C++
PyTorch C++
Databases
PostgreSQL
Weaviate
Milvus
pgvector
Pinecone
Apache Kafka
AI/ML
Weights & Biases
LangChain
Spark
LoRA
Fine-tuning
Embeddings
JAX
Prompt Engineering
PEFT
Transformers
TensorFlow
PyTorch
LLM
RAG
Ray
Hugging Face
Machine Learning
Mobile
JUnit
DevOps
Rest API
gRPC
GitHub Actions
OpenTelemetry
Prometheus
GitLab CI
CI/CD
Git
Docker
Kubernetes
Grafana
Analytics
ETL/ELT
QA
Pytest
Apply
$106k – $158k per year • Equity • Remote/Hybrid • Full-Time • Bachelor's Degree • Seattle
Python
Go
Java
C++
Python
Flask
FastAPI
C++
TensorFlow C++
PyTorch C++
Databases
PostgreSQL
Weaviate
Milvus
pgvector
Pinecone
Apache Kafka
AI/ML
Weights & Biases
LangChain
Spark
LoRA
Fine-tuning
Embeddings
JAX
Prompt Engineering
PEFT
Transformers
TensorFlow
PyTorch
LLM
RAG
Ray
Hugging Face
Machine Learning
Mobile
JUnit
DevOps
Rest API
gRPC
GitHub Actions
OpenTelemetry
Prometheus
GitLab CI
CI/CD
Git
Docker
Kubernetes
Grafana
Analytics
ETL/ELT
QA
Pytest
Apply
$59k – $137k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Mexico City
Python
Python
Asyncio
Databases
Apache Kafka
AI/ML
AI Agents
LLM
OCR
DevOps
KEDA
CI/CD
AWS
Kubernetes
AWS Step Functions
Apply
$81k – $210k per year (Estimated) • In office • Full-Time • Waterloo • Toronto
Python
PowerShell
AI/ML
Copilot
OpenAI
DevOps
Terraform
Azure
CI/CD
Kubernetes
Bicep
Azure AKS
Apply
$38k – $100k per year (Estimated) • Remote/Hybrid • Full-Time • São Paulo
Python
AI/ML
Machine Learning
Apply
$350k – $425k per year • Remote • Full-Time • 12+ years exp • Denver • New York • San Francisco
AI/ML
Runway
World Models
Apply
$150k – $200k per year • In office • Full-Time • 7+ years exp • New York
AI/ML
Runway
World Models
Apply
$130k – $175k per year • Remote • Full-Time • 5+ years exp • New York
AI/ML
Runway
World Models
Management
Agile
Apply
$240k – $290k per year • Remote • Full-Time • 5+ years exp
JavaScript
TypeScript
Node JS
AI/ML
Computer Vision
LLM
Runway
World Models
Machine Learning
Frontend
React.js
DevOps
AWS
AWS Fargate
Amazon ECS
Apply
GTM Engineer 22 days ago
$175k – $225k per year • Remote • Full-Time • 6+ years exp • New York • Seattle • San Francisco
Python
SQL
AI/ML
Runway
LLM Guardrails
Agentic Workflows
World Models
Marketing
Salesforce
Apply
See all jobs
This is one of many
713,007 more open roles from verified company boards, updated every day.