368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$143k – $268k per year (Estimated)
Location
Remote/Hybrid (Palo Alto, United States)
Seniority
Middle · 4+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Mistral AI is a French artificial intelligence company founded in Paris in 2023 by former researchers from Google DeepMind and Meta. It builds efficient open-weight and commercial large language models, including the Mistral, Mixtral, Magistral and Codestral families, and distributes them through its own developer platform as well as the major cloud marketplaces. Positioned as Europe's leading frontier model developer, the company also ships the Le Chat assistant, on-premise deployment options for regulated industries and a sovereign compute offering for customers who must keep data inside the region.

About Mistral

Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems-across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector-co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.

The Role

This role focuses on building and operating the end-to-end execution, training, and data infrastructure that powers Mistral’s agentic models and coding assistants. You will be a core contributor to our agent research stack: designing scalable systems for synthetic data generation, building ultra-fast training and RL execution environments, and maintaining high-throughput execution engines.

You will tackle the engineering challenges at every step of the agent lifecycle: from orchestrating 1M+ concurrent and short-lived sandboxes for untrusted code execution to optimizing agent training codebases, distributed trajectory collection pipelines, and dataset processing workflows across massive hybrid and multi-cloud clusters.

What You Will Do

  • Large-Scale Sandboxing Infrastructure: Design, deploy, and operate our high-throughput sandboxing platform, executing LLM-generated untrusted code across over 1 million isolated environments concurrently for model evaluation and interactive RL environments.

  • Agent Data Generation Pipelines: Architect and scale high-throughput pipelines for synthetic code generation, agent trajectories, rollouts, and self-play data collection to power post-training and RL loops.

  • Training Codebase & Systems Optimization: Optimize agent training codebases and distributed execution runtimes (PyTorch, Ray, SLURM/Kubernetes) to minimize multi-step rollout overhead, improve GPU utilization, and eliminate scaling bottlenecks.

  • Low-Latency Orchestration & Warm Pooling: Reduce sandbox cold-start times to sub-second levels using container warm pools, snapshot/restore technology (e.g., CRIU, microVMs), and optimized image delivery layers across hybrid clusters.

  • Multi-Cluster Queueing & Resource Allocation: Implement Kubernetes-native custom controllers, CRDs, and queuing systems to dynamically route short-lived evaluation, synthetic data, and agent execution tasks across diverse hardware fleets.

  • Isolation, Security & Security Boundary: Ensure strict multi-tenant network and process isolation for untrusted agent code using container/sandboxing runtimes (e.g., gVisor, Firecracker) and default-deny network postures.

  • Operational Excellence: Maintain high availability, telemetry, and automated self-healing across millions of transient jobs while participating in on-call rotations for critical agent training and execution pipelines.

What We're Looking For

  • 4+ years of experience in Systems Engineering, Distributed Systems, Cloud Infrastructure, or MLOps supporting LLM/RL workloads.

  • Data & Pipeline Engineering: Proven experience building high-throughput data processing and generation pipelines for large-scale datasets (e.g., Ray, Spark, custom distributed queues).

  • Deep experience with Kubernetes & Container Tech: Strong expertise writing custom K8s operators/controllers, managing Linux cgroups/namespaces, and optimizing Docker image layers and distribution systems.

  • High-Performance Software Engineering: Advanced proficiency in Python, Go, C++ or Rust, with a track record of profiling and optimizing high-performance ML or backend systems codebases.

  • Sandboxing & Isolation Technologies: Hands-on experience with lightweight virtualization, container runtimes, or WASM (e.g., Docker, gVisor, Firecracker).

  • Queueing & Scheduling: Deep familiarity with task queue systems, resource schedulers, and low-latency queuing architectures for high-volume, short-lived workloads.

  • Comfort with Ambiguity: Passion for working directly alongside AI researchers to rapidly turn frontier agent ideas into scalable, production-grade infrastructure.

What We Offer

We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.

For the most up-to-date details on benefits available in your location, please refer to our Benefits page.

Privacy Policy

Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Palo Alto
$136k – $253k per year • Equity • Remote/Hybrid • Full-Time • 10+ years exp • Frisco • New York • Toronto • Ann Arbor
Python
SQL
Java
Java
Flyway
AI/ML
AWS Bedrock
Claude
LLM
Anthropic
AWS Bedrock AgentCore
LLM Guardrails
LLMOps
DevOps
Amazon EKS
AWS
CI/CD
Datadog
Docker
Kubernetes
Apply
$73k – $183k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Zug
Python
Python
FastAPI
Databases
PostgreSQL
Redis
AI/ML
AI Agents
AWS Bedrock
AWS Bedrock AgentCore
LLM
DevOps
Amazon EKS
AWS
AWS CDK
CI/CD
Datadog
Kubernetes
OpenTelemetry
Platform Engineering
Apply
SR AI ENGINEER, SMAI 6 hours ago
In office • Full-Time • 2+ years exp • Bachelor's Degree • Taoyuan
JavaScript
Python
SQL
TypeScript
Python
FastAPI
AI/ML
AI Agents
Function Calling
LLMOps
Prompt Engineering
Quantization
RAG
Streamlit
Frontend
Angular
React.js
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
GitHub Actions
Kubernetes
OpenShift
Apply
$32k – $42k per year • Remote/Hybrid • Full-Time • 4+ years exp • Cyprus
Java
SQL
Java
Spring Boot
Databases
Apache Kafka
Oracle
DevOps
AWS
Docker
Jenkins
Kubernetes
Rest API
Apply
$24k – $63k per year (Estimated) • In office • Full-Time • 6+ years exp • Master's Degree • India
Crystal
Groovy
JavaScript
Perl
Python
Ruby
SQL
TypeScript
Java
Java
Apache Tomcat
Gradle
Hibernate
Maven
Spring Boot
Spring MVC
Databases
Apache Kafka
Db2
Oracle
PostgreSQL
RabbitMQ
AI/ML
Fine-tuning
Frontend
Angular
JQuery
DevOps
Apache HTTP Server
AWS
Azure
CI/CD
Docker
GCP
Jenkins
Kubernetes
Rest API
Cybersecurity
Checkmarx
SonarQube
Apply
$77k – $184k per year (Estimated) • In office • Full-Time • Master's Degree • Paris • Amsterdam • Linz • Munich • London
Python
AI/ML
JAX
DevOps
HPC
Apply
$76k – $180k per year (Estimated) • Remote/Hybrid • Full-Time • Paris • Amsterdam • Linz • Munich • London
Python
AI/ML
Fine-tuning
LLM
Mistral
RLHF
Post-training
SFT
AI Agents
Function Calling
DevOps
HPC
Design
SolidWorks
Apply
In office • Full-Time • Paris
DevOps
AWS
Azure
GCP
IAM
Cybersecurity
GDPR
HIPAA
ISO 27001
NIST CSF
SOC 2
Apply
$54k – $103k per year (Estimated) • In office • Full-Time • Paris • Amsterdam • Warsaw • Munich • London
Go
DevOps
Ansible
Kubernetes
Terraform
Cybersecurity
FortiGate
Apply
$42k – $123k per year (Estimated) • In office • Full-Time • 2+ years exp • Paris
Python
JavaScript
Python
FastAPI
AI/ML
RAG
Frontend
React.js
Vue.js
Apply
$110k – $240k per year (Estimated) • In office • Bachelor's Degree • Palo Alto
Java
Python
Scala
AI/ML
AI Agents
Fine-tuning
LLM Guardrails
DevOps
AWS
Azure
GCP
Git
GitHub
Marketing
Salesforce
Apply
Chief of Staff 9 hours ago
$120k – $150k per year • Equity 0.4–0.7% • In office • Full-Time • 3+ years exp • Palo Alto
AI/ML
AI Agents
Apply
$140k – $310k per year (Estimated) • Remote/Hybrid • Bachelor's Degree • Palo Alto
Databases
Apache Kafka
NATS
DevOps
AWS
Azure
CI/CD
Docker
GCP
Grafana
gRPC
Kubernetes
OpenTelemetry
Platform Engineering
Prometheus
Robotics
EtherCAT
IoT
MQTT
OPC UA
Apply
$139k – $294k per year (Estimated) • In office • Palo Alto
Python
AI/ML
Fine-tuning
Hybrid Search
LLM
Prompt Engineering
RAG
Human-in-the-Loop
Knowledge Graph
AI Agents
Function Calling
DevOps
AWS
Apply
$137k – $292k per year (Estimated) • In office • Palo Alto
Python
AI/ML
Hybrid Search
LLM
Prompt Engineering
RAG
Human-in-the-Loop
AI Agents
Function Calling
DevOps
AWS
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.