386,639open jobs
10,098companies
50,285added this week
Browse all
Salary
$95k – $200k per year (Estimated)
Location
Remote/Hybrid (Palo Alto, United States)
Seniority
Middle · 4+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Mistral AI is a French artificial intelligence company founded in Paris in 2023 by former researchers from Google DeepMind and Meta. It builds efficient open-weight and commercial large language models, including the Mistral, Mixtral, Magistral and Codestral families, and distributes them through its own developer platform as well as the major cloud marketplaces. Positioned as Europe's leading frontier model developer, the company also ships the Le Chat assistant, on-premise deployment options for regulated industries and a sovereign compute offering for customers who must keep data inside the region.

About Mistral

Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems-across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector-co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.

The Role

This role focuses on building and operating the ML platform that powers large-scale training, evaluation, and batch inference at Mistral AI. You will develop the infrastructure that enables researchers and engineers to run distributed GPU workloads reliably across clusters, hardware types, and regions.

You will work across the full ML lifecycle, from workload scheduling and capacity management to platform APIs, observability, and production operations. You will take ownership of critical systems and help turn complex infrastructure into reliable, self-service capabilities.

What You Will Do

  • Build the ML Platform: Develop services, APIs, controllers, and tooling for training, evaluation, fine-tuning, and batch inference.

  • Orchestrate GPU Workloads: Build systems for queueing, admission control, quotas, priorities, preemption, and topology-aware placement.

  • Manage Compute Capacity: Improve how heterogeneous GPU resources are provisioned, allocated, and utilized across clusters.

  • Enable Multi-Cluster Execution: Place workloads based on capacity, data locality, hardware requirements, and organizational priorities.

  • Improve Researcher Experience: Create self-service workflows that make distributed workloads easy to launch, observe, debug, and reproduce.

  • Optimize Performance: Improve GPU utilization, scheduling latency, workload startup time, throughput, and infrastructure efficiency.

  • Build for Reliability: Develop observability, failure recovery, capacity planning, and operational tooling for critical ML workloads.

  • Operate What You Build: Participate in on-call rotations and troubleshoot issues across applications, schedulers, networking, storage, and GPU infrastructure.

What We're Looking For

  • Have 4+ years of experience in ML infrastructure, distributed systems, Kubernetes platform engineering, or a related field.

  • Are proficient in Python or Go and comfortable working with production-grade distributed systems.

  • Have strong Kubernetes knowledge, including controllers, operators, CRDs, scheduling, networking, storage, and resource management.

  • Understand technologies such as Kueue, Karpenter, Volcano, and Kyverno, and the problems they address in workload scheduling, provisioning, and policy enforcement.

  • Understand distributed ML workloads, including training, fine-tuning, evaluation, checkpointing, and batch inference.

  • Are familiar with GPU infrastructure and technologies such as PyTorch, CUDA, NCCL, and high-performance networking.

  • Understand concepts such as quotas, priorities, preemption, gang scheduling, topology awareness, and workload admission.

  • Can diagnose performance and reliability problems across software, orchestration, networking, storage, and hardware.

  • Care about developer experience and enjoy turning complex infrastructure into simple, reliable interfaces.

  • Thrive in an ambiguous, fast-moving environment shaped by frontier AI research.

What We Offer

We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.

For the most up-to-date details on benefits available in your location, please refer to our Benefits page.

Privacy Policy

Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
386,639 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Palo Alto
$34k – $81k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Bengaluru • Chennai
Python
AI/ML
AI Agents
Embeddings
Fine-tuning
NLP
RAG
Sentiment Analysis
Speech Recognition
Apply
Senior Data Scientist 2 hours ago
$32k – $63k per year (Estimated) • In office • PhD • Bengaluru
AI/ML
AWQ
Bitsandbytes
Fine-tuning
GGUF
LLM
LoRA
NLP
PEFT
Prompt Engineering
QLoRA
Transformers
Apply
AI Engineer 2 hours ago
$169k – $316k per year (Estimated) • Remote • Internship • 10+ years exp • Master's Degree
AI/ML
Fine-tuning
LLM
Post-training
Pre-training
Cybersecurity
Defense in Depth
Least Privilege
Zero Trust
Apply
SDE II - Fullstack 2 hours ago
$20k – $53k per year (Estimated) • In office • 3+ years exp • Noida
Go
JavaScript
Node JS
PHP
Python
TypeScript
Databases
MySQL
Frontend
Angular
React.js
Vue.js
DevOps
AWS
CI/CD
Apply
$15k – $37k per year (Estimated) • Remote • 2+ years exp • Moscow
Go
Python
Cybersecurity
Acunetix
Nessus
Apply
$58k – $154k per year (Estimated) • In office • Full-Time • Paris
C#
Go
Python
TypeScript
Python
Django
FastAPI
Flask
Apply
$77k – $184k per year (Estimated) • In office • Full-Time • Master's Degree • Paris • Amsterdam • Linz • Munich • London
Python
AI/ML
JAX
DevOps
HPC
Apply
$76k – $180k per year (Estimated) • Remote/Hybrid • Full-Time • Paris • Amsterdam • Linz • Munich • London
Python
AI/ML
Fine-tuning
LLM
Mistral
RLHF
Post-training
SFT
AI Agents
Function Calling
DevOps
HPC
Design
SolidWorks
Apply
In office • Full-Time • Paris
DevOps
AWS
Azure
GCP
IAM
Cybersecurity
GDPR
HIPAA
ISO 27001
NIST CSF
SOC 2
Apply
$54k – $103k per year (Estimated) • In office • Full-Time • Paris • Amsterdam • Warsaw • Munich • London
Go
DevOps
Ansible
Kubernetes
Terraform
Cybersecurity
FortiGate
Apply
$95k – $193k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Palo Alto
Python
SQL
Databases
Databricks
Google BigQuery
Snowflake
BigQuery
AI/ML
LLM
NLP
Sentiment Analysis
DevOps
SLI/SLO/SLA
Analytics
Power BI
Tableau
Marketing
YouTube
Apply
$179k – $269k per year • Equity • Remote • Full-Time • 5+ years exp • Bachelor's Degree • Palo Alto
C++
Python
C++
PyTorch C++
TensorFlow C++
AI/ML
Computer Vision
PyTorch
TensorFlow
Robotics
Motion Planning
Sensor Fusion
Apply
$87k – $173k per year (Estimated) • Remote • Contractor • 6+ years exp • Palo Alto
Apply
$157k – $286k per year (Estimated) • Equity • Remote/Hybrid • 12+ years exp • Bachelor's Degree • Palo Alto
DevOps
AWS
Azure
GCP
Marketing
Salesforce
Apply
$100k – $243k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Palo Alto
TypeScript
JavaScript
AI/ML
AI Agents
Function Calling
Human-in-the-Loop
LLM
Prompt Engineering
RAG
Frontend
React.js
Apply
See all jobs
This is one of many
386,639 more open roles from verified company boards, updated every day.