585,223open jobs
25,872companies
81,805added this week
Browse all
Salary
$68k – $182k per year (Estimated)
Location
In office (Berlin, Freiburg im Breisgau, New York)
Seniority
Middle · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match

Who we are

Foundation models transformed text and images. Structured data - the largest and most consequential data format in the world - stayed untouched, until now. What LLMs did for language, we're doing for tables.

We pioneered tabular foundation models: TabPFN v2 was a Nature cover story, has passed 3.5M+ downloads and 7,500+ GitHub stars, and runs in production from detecting lung disease with Oxford Cancer Analytics to preventing train failures with Hitachi. The hardest problems - millions of rows, real-time inference, entirely new modalities - are still open, and no one else is working on them at this level.

We're a small, highly selective team of 40+ with backgrounds from Google, DeepMind, Meta, Apple, Amazon, Jane Street, and CERN, led by Frank Hutter, Noah Hollmann, and Sauraj Gambhir, and advised by Bernhard Schölkopf and Turing Award winner Yann LeCun.

In July 2026, less than 18 months after our €9M pre-seed, we joined SAP as an independent frontier AI lab - same team, mission, and open-weights models, now backed by more than €1 billion over four years.

About the Role

We spend tens of millions per year on GPU compute to train tabular foundation models. That's not a target, it's what we're running today, and it's growing. The person who owns this infrastructure makes decisions worth millions of dollars: cluster architecture, scheduling efficiency, provider strategy, hardware selection. A wrong call costs six figures.

Today we run Slurm on GCP across multiple clusters. We're scaling to multi-cluster, multi-provider infrastructure and evaluating new hardware generations as they come online. You own the full stack, from cluster operations and cost optimization to distributed training performance and the tooling layer that keeps researchers moving fast. You work directly with the research team and understand what they're doing well enough to make infrastructure decisions that actually help them. And this isn't a pure support role. We operate an open environment. If you've got the next SOTA tabular architecture up your sleeve, go ahead and train it.

What you'll work on:

  • Own and evolve multi-cluster GPU infrastructure. Slurm on GCP today, multi-provider and new hardware tomorrow. Architecture, scheduling, reliability, cost optimization

  • Drive GPU utilization and training throughput: profiling, memory optimization, communication bottlenecks, systems-level debugging of distributed training across large runs

  • Architect the next generation of our infrastructure: multi-cluster orchestration, new GPU generations, provider diversification, capacity planning against growing compute demands

  • Build the developer productivity layer: CI pipelines, experiment tracking, model registry, data processing, and internal tooling that keeps research iteration speed high

  • Own the compute budget. You understand cost per FLOP across providers and hardware, and you hate wasted compute

Tech stack: Slurm, GCP, Docker, wandb, GitHub Actions, uv, PyTorch, Triton

You may be a good fit if you have:

  • 3+ years building and operating production GPU infrastructure or distributed training systems at scale. At a major AI lab, a well-funded ML startup, or an HPC environment

  • Deep hands-on experience with Slurm and cluster management. You've debugged scheduling failures, optimized utilization across multi-tenant GPU workloads, and operated infrastructure where downtime has real cost

  • Expert-level systems thinking: memory bandwidth, GPU profiling. You reason about hardware, not configs

  • Strong Python and genuine fluency with PyTorch internals. Enough to profile a training run and tell whether the bottleneck is data loading, communication, or compute

  • Track record of making infrastructure decisions that measurably improved training throughput or cost efficiency

  • Strong AI tooling skills. You use Claude Code, Cursor, or similar fluently to move fast without sacrificing quality

Bonus:

  • Experience operating at tens-of-millions-scale GPU spend

  • Multi-cloud or hybrid HPC/cloud infrastructure experience

  • Triton, CUDA, or custom kernel experience

  • Experience scaling from single cluster to multi-cluster orchestration

  • Background building experiment tracking, model registry, or ML pipeline tooling

Life at Prior Labs

You'll work alongside researchers and builders who hold themselves to a very high bar - in the quality of their work and in how they work with each other. We move fast and still take the time to do things right.

Our teams are based in Berlin, Freiburg, and New York - when you're working on something as hard as TabPFN, being in the same room matters. But great people come from everywhere, and in exceptional cases we're open to remote, which usually means frequent travel to one of our offices. Wherever you're based, the whole company comes together regularly for offsites to build and celebrate together.

Our Commitments

The best products and teams are built by people with a wide range of perspectives and backgrounds. We welcome applications from all identities and walks of life - especially if you've ever felt discouraged by "not checking every box" - and provide equal opportunities regardless of gender, sexual orientation, origin, disability, or any other trait that makes you who you are.

We care about how your data is handled - see our Recruiting Data Privacy page

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
585,223 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Berlin
$87k – $157k per year • Remote/Hybrid • Public Trust • Full-Time • 4+ years exp • Bachelor's Degree • Gaithersburg
Python
Java
C++
Databases
Apache Kafka
AI/ML
Claude Code
DevOps
GCP
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Harbor
Management
Agile
Apply
$87k – $157k per year • Remote/Hybrid • Public Trust • Full-Time • 4+ years exp • Bachelor's Degree • United States
Python
Java
C++
Databases
Apache Kafka
AI/ML
Claude Code
DevOps
GCP
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Harbor
Management
Agile
Apply
$117k – $133k per year • In office • Part-Time • 1+ year exp • Bachelor's Degree • Cambridge
Python
AI/ML
Claude Code
Apply
$140k – $273k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • McLean
Python
Java
SQL
Scala
Databases
MySQL
MongoDB
Snowflake
Cassandra
Apache Kafka
Amazon Redshift
AI/ML
Hadoop
Spark
DevOps
GCP
Azure
AWS
Management
Agile
Apply
$209k – $239k per year • In office • Full-Time • 9+ years exp • Bachelor's Degree • Wilmington
Python
Java
SQL
Scala
Databases
MySQL
MongoDB
Snowflake
Cassandra
Apache Kafka
Amazon Redshift
AI/ML
Hadoop
Spark
DevOps
GCP
Azure
AWS
Management
Agile
Apply
$68k – $136k per year (Estimated) • In office • Full-Time • 2+ years exp • Berlin
Databases
Databricks
DevOps
GCP
Azure
AWS
GitHub
Apply
$120k – $239k per year (Estimated) • In office • Full-Time • Berlin
DevOps
GitHub
Cybersecurity
ISO 27001
SOC 2
GDPR
Apply
In office • Full-Time • New York • Berlin
DevOps
GitHub
Apply
In office • Full-Time • New York • Berlin
DevOps
GitHub
Apply
$50k – $125k per year (Estimated) • In office • Full-Time • Berlin
AI/ML
OpenAI
Anthropic
DevOps
GitHub
Apply
$72k per year • Remote/Hybrid • Full-Time • Berlin • Hanover
Python
Apply
$70k – $140k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Berlin
JavaScript
Node JS
Node JS
Nest.JS
AI/ML
Function Calling
AI Agents
LLM
RAG
Tool Use
Apply
Data Scientist 6 hours ago
$82k – $128k per year • In office • Full-Time • 10+ years exp • PhD • Berlin
AI/ML
Cursor
Claude Code
AI Agents
Apply
$68k – $162k per year (Estimated) • Remote/Hybrid • Berlin
Apply
$70k – $82k per year • Remote/Hybrid • Full-Time • 5+ years exp • Berlin
Apply
See all jobs
This is one of many
585,223 more open roles from verified company boards, updated every day.