712,228open jobs
42,462companies
99,345added this week
Browse all
Salary
$71k – $189k per year (Estimated)
Location
In office (London, Munich)
Seniority
Middle · 3+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 23, 2026. First seen by Alion on May 6, 2026. SpAItial scores C on the Alion truth index.

Overview
Company
Impact
Profile match
SpAItial is building Echo, a frontier world model that turns images and prompts into persistent, explorable 3D Gaussian Splat worlds.

SpAItial is pioneering the next generation of World Models, pushing the boundaries of generative AI, computer vision, and simulation. We are moving beyond 2D pixels to build models that natively understand the physics and geometry of our world. Our mission is to redefine how industries, from robotics and AR/VR to gaming and cinema, generate and interact with physically-grounded 3D environments.

We’re looking for bold, innovative individuals driven by a passion for tackling hard problems in generative 3D AI. You should thrive in an environment where creativity meets technical challenge, take pride in craft, and collaborate closely with a small team building frontier systems.

We are seeking a Machine Learning Systems & Infrastructure Engineer to build and own the systems that turn raw real-world data into trained world models and reliable production endpoints. You will design, implement, and operate scalable training stacks, data ingestion pipelines, experiment orchestration, and model serving for large diffusion-based generative models. The role is hands-on and code-heavy - you will work inside the same monorepo as the research team, mostly in Python, and should be as comfortable refactoring a trainer class or a dataset loader as you are writing Terraform.

Responsibilities

  • Own and evolve the ML systems that enable training, evaluation, and serving of large foundation models - trainer, dataset loaders, checkpointing, and experiment orchestration code.

  • Distributed training enablement: Improve high-throughput training stacks (e.g., PyTorch DDP/FSDP, NCCL) for performance, stability, and reproducibility, including preemption-safe and sharded checkpointing.

  • Data systems and pipelines: Build end-to-end Python pipelines that turn third-party capture sources into clean, versioned training datasets - including scraping (e.g., Playwright) and preprocessing - and optimize the underlying storage at petabyte scale (object storage, fuse mounts, caching layers, shared filesystems, and relational / analytical / embedded metadata stores).

  • ML workflow orchestration and serving: Operate the systems researchers use to launch experiments, data jobs, and production endpoints - workflow engines (e.g., Kubeflow Pipelines, Airflow), GPU schedulers (e.g., Volcano, Slurm), experiment trackers (e.g., MLflow, Weights & Biases), and managed-inference platforms (e.g., Modal, Triton) - and maintain a launcher SDK for one-command runs.

  • Containerization and packaging: Ship workloads with Docker and Kubernetes; maintain IaC (Terraform) for the surfaces you own and CI/CD pipelines, including self-hosted GPU runners.

  • Observability and reliability: Monitoring, logging, and alerting for job performance, data-pipeline health, and cost (e.g., Prometheus/Grafana, OpenTelemetry); define SLOs and incident response for the systems you own.

  • Security and access: Manage secrets, IAM, and network boundaries (e.g., Tailscale, cloud VPC) for the systems you own.

  • Collaboration: Partner with ML researchers, engineers, and the platform team to unblock training and data work and improve developer experience.

Key Qualifications

  • 3+ years writing production-quality Python in a large, multi-author codebase, with strong SWE fundamentals (ML systems experience strongly preferred).

  • Hands-on with modern ML training stacks (PyTorch; DDP/FSDP or comparable); have personally debugged distributed jobs across many GPUs and nodes.

  • Have shipped non-trivial end-to-end data pipelines at scale - ingestion, transformation, validation, versioning, republish - ideally including real-world sources with rate limits, auth, or undocumented APIs.

  • Hands-on GPU compute and performance debugging (CUDA/NCCL, GPU utilization, networking bottlenecks, profiling).

  • Working knowledge of cloud environments (AWS, GCP, or Azure), including object storage, IAM, and cost awareness.

  • Proficient with containers (Docker, Kubernetes) and comfortable reading and writing IaC (Terraform) for the surfaces you ship.

  • Strong working knowledge of how to store and query large datasets at scale: SQL fundamentals; relational (e.g., Postgres), analytical (e.g., BigQuery, Snowflake), and embedded (e.g., SQLite) stores; and object storage with caching layers. Familiarity with ML workflow orchestration and experiment tracking (e.g., Kubeflow Pipelines, MLflow).

  • Experience with monitoring and observability tooling (e.g., Prometheus/Grafana, OpenTelemetry) and CI/CD for infra and ML workflows (e.g., GitHub Actions).

At SpAItial, we are committed to creating a diverse and inclusive workplace. We welcome applications from people of all backgrounds, experiences, and perspectives. We are an equal opportunity employer and ensure all candidates are treated fairly throughout the recruitment process.

At SpAItial, we are committed to creating a diverse and inclusive workplace. We welcome applications from people of all backgrounds, experiences, and perspectives. We are an equal opportunity employer and ensure all candidates are treated fairly throughout the recruitment process.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
712,228 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
London
$128k – $180k per year • In office • Full-Time • 3+ years exp • New York
Python
Java
SQL
DevOps
Incident Management
Apply
$100k – $130k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • Atlanta
Python
Go
TypeScript
SQL
AI/ML
Ray
DevOps
Terraform
GitHub Actions
CloudFormation
CI/CD
Jenkins
AWS
Docker
Kubernetes
Amazon EKS
AWS Fargate
AWS Lambda
Amazon EC2
GitLab
Amazon S3
IAM
Amazon ECS
Amazon CloudWatch
Apply
$70k – $150k per year • In office • Full-Time • 5+ years exp • Toronto
Python
SQL
Analytics
Power BI
Apply
$97k – $170k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Toronto
SQL
DevOps
Rest API
Azure
AWS
Analytics
Power BI
ETL/ELT
SSIS
Management
Power Automate
Power Apps
SharePoint
Agile
Apply
$117k – $338k per year (Estimated) • Remote • Internship • New York • Denver
SQL
Apply
Technical Artist 30 days ago
In office • Full-Time • 3+ years exp • London • Munich
JavaScript
AI/ML
World Models
Frontend
Three.JS
Game Dev
Houdini
Design
Blender
Adobe After Effects
Apply
$57k – $143k per year (Estimated) • In office • Full-Time • London
AI/ML
Computer Vision
World Models
Machine Learning
Management
Notion
Google Workspace
Xero
Apply
$102k – $265k per year (Estimated) • In office • Full-Time • PhD • London • Munich
Python
AI/ML
Fine-tuning
Computer Vision
VLM
PyTorch
Tokenization
SFT
Post-training
FSDP
Vision-Language-Action
World Models
Machine Learning
Robotics
Sim-to-Real
Imitation Learning
Apply
$102k – $265k per year (Estimated) • In office • Full-Time • PhD • London • Munich
Python
AI/ML
Computer Vision
PyTorch
World Models
Robotics
ORB-SLAM3
COLMAP
Apply
$72k – $192k per year (Estimated) • In office • Full-Time • 3+ years exp • London • Munich
Python
PowerShell
AI/ML
CUDA Toolkit
Computer Vision
PyTorch
CUDA
FSDP
NCCL
World Models
Machine Learning
DevOps
Terraform
GCP
GitHub Actions
OpenTelemetry
CircleCI
Prometheus
Azure
CI/CD
AWS
Docker
Kubernetes
Grafana
IAM
Apply
CloudOps Lead 1 day ago
$104k – $226k per year (Estimated) • In office • Full-Time • 10+ years exp • London
Python
DevOps
Terraform
Ansible
CI/CD
AWS
Linux
Apply
$66k per year • Remote/Hybrid • Full-Time • London
Management
Google Workspace
Apply
In office • Contractor • London
Management
Microsoft Office
Apply
$41k – $110k per year (Estimated) • Remote/Hybrid • Full-Time • London
Apply
$37k – $62k per year (Estimated) • In office • Full-Time • London
AI/ML
AI Agents
Apply
See all jobs
This is one of many
712,228 more open roles from verified company boards, updated every day.