368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$71k – $171k per year (Estimated)
Location
Remote/Hybrid (Paris, France)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Our true mission is to bring you the cloud that makes sense. Scaleway is the obvious scale partner that all developers, architects and business owners are looking for: in short, the cloud that makes sense. We are fulfilling our true mission of ...

WHY WE NEED YOU ?

As our GPU Cloud infrastructure continues to scale, we are strengthening our SRE organization to support the deployment and operation of increasingly large and complex AI and HPC infrastructure.

Your mission will be to lead our Site Reliability Engineering team and ensure the reliability, scalability, and operational excellence of our GPU clusters.

This is not a traditional IT Operations management role. You will combine engineering leadership with strong technical ownership, working close to the infrastructure itself from Linux systems, networking and bare-metal server provisioning to hardware lifecycle, automation and cluster observability.

You will help the team automate critical infrastructure workflows, improve reliability and operate production-grade GPU platforms powering our sovereign cloud.

YOUR FUTURE TEAM

We work in a collaborative and international environment where the diversity of Scalers, combined with a strong culture of knowledge sharing, helps us bring ambitious projects to life.

You will lead a team of 6 Site Reliability Engineers within the GPU Cloud organization.

The team works on some of our most critical AI and HPC infrastructure challenges, including bare-metal provisioning, GPU cluster automation, server lifecycle management, hardware failure management, observability, reliability and the integration of new GPU technologies.

The scope goes beyond traditional cloud-native infrastructure: the team operates close to the physical servers and needs to automate the full lifecycle of large fleets of GPU machines, from remote provisioning to production operations and remediation.

You will collaborate closely with GPU Cloud Engineering, Hardware, Product and Operations teams, as well as other infrastructure teams across Scaleway.

YOUR DAILY ROUTINE

Tasks

  • Lead and manage a team of 6 Site Reliability Engineers, supporting both their technical execution and career development
  • Provide technical leadership and challenge architecture decisions related to large-scale GPU infrastructure
  • Design and drive automation for bare-metal server provisioning and lifecycle management
  • Improve remote deployment and server management capabilities using technologies and concepts such as PXE, BMC and IPMI
  • Drive automation around hardware failures, remediation and server recovery
  • Build and improve observability, monitoring, logging and alerting capabilities across production GPU clusters
  • Own the SRE team's technical roadmap, priorities and delivery
  • Ensure the reliability, scalability, performance and resilience of production GPU infrastructure
  • Drive continuous improvements in automation, incident response and post-incident remediation
  • Support technical decisions involving Linux systems, networking, hardware and cluster architecture
  • Collaborate closely with Engineering, Hardware, Product, Operations and other GPU Cloud teams
  • Recruit, onboard, coach and develop engineers within the team
  • Lead incident response and post-incident improvements when critical production issues occur

ABOUT YOU

HARDSKILLS:

  • Strong experience managing SRE, Infrastructure or Platform Engineering teams in production environments
  • Strong knowledge of Linux/UNIX systems and infrastructure fundamentals
  • Good networking knowledge, including VLAN, VRF, VIP and NAT
  • Experience with bare-metal infrastructure and remote server provisioning using technologies such as PXE, BMC or IPMI
  • Strong understanding of infrastructure automation, observability and incident management
  • Experience with Kubernetes and monitoring stacks such as Prometheus and Grafana
  • Experience with GPU, HPC or high-performance infrastructure is a strong plus

SOFT SKILLS:

  • Strong engineering leadership and team management capabilities

  • Technical rigor and high attention to detail in production-critical environments

  • Ability to handle high-pressure operational situations and manage incident stress pragmatically

  • Excellent communication skills with the ability to convey challenging messages effectively

  • Collaborative mindset with a focus on empowering engineers rather than micromanaging

WHAT YOU WILL FIND AT SCALEWAY ++++

  • Hybrid work: We offer up to 3 days of remote work per week.

  • Offices: Our offices are spacious, dynamic workspaces with bold design, conveniently located near public transport. Most of our offices feature outdoor spaces (terraces) and bike parking facilities.

  • Dining: Our chef provides a healthy meal service at the headquarters, and breakfast is available across all our sites year-round. Scalers working from regional sites enjoy a Swile card for lunches.

  • Well-being commitments: Whether it’s access to a gym, daycare places, or discounted services for caring services, Scaleway is committed to supporting Scalers in maintaining a balanced life.

  • International environment: With dozens of nationalities, Scaleway offers a stimulating environment where English is as widely spoken as French.

  • Career & Mobility: Our managers value internal mobility, and opportunities to transition to other entities within the Iliad Group are accessible to all Scalers.

Why join the Scaleway adventure?

A rich and diverse product offering: Scaleway offers over 100 public cloud products in IaaS, PaaS, and AI.

A cutting-edge technical environment: Scaleway provides modern infrastructures, including high-performance bare metal servers, to tackle exciting technical challenges.

Commitment to responsible cloud: Scaleway is dedicated to a more responsible cloud, with data centers powered solely by renewable energy since 2017, minimizing our ecological footprint and holding top-level certification.

THE NEXT STEPS …

  • Discovery call with HR
  • Technical interview with the HPC team to understand your technical skills and approach to the role
  • Manager interview to validate your expertise
  • Interview with an Engineering Manager / Head of Engineering to deepen discussions and assess your fit with the team
  • HR interview and office visit to tour our offices and meet your future colleagues
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Paris
$25k – $42k per year • Equity 0–0.2% • Remote • Full-Time • 3+ years exp
Bash
Go
JavaScript
Python
TypeScript
DevOps
AWS
Azure
CI/CD
Datadog
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Incident Management
Kubernetes
Platform Engineering
Prometheus
Terraform
Amazon CloudWatch
GitHub
GitLab
IAM
Cybersecurity
Least Privilege
Apply
Founding Engineer 1 day ago
$81k – $116k per year • In office • Full-Time • Bachelor's Degree • Munich
JavaScript
Python
TypeScript
Databases
MySQL
PostgreSQL
Frontend
Next.js
React.js
Tailwind CSS
DevOps
AWS
Azure
CI/CD
Docker
GCP
Grafana
Kubernetes
OpenTelemetry
Prometheus
Apply
$84k – $178k per year (Estimated) • In office • Full-Time • 10+ years exp • Wellington
Java
Python
DevOps
Ansible
AWS
Azure
CI/CD
Docker
GCP
Helm
Kubernetes
Platform Engineering
Prometheus
Service Mesh
Terraform
GitLab
IAM
Apply
$98k – $195k per year (Estimated) • In office • Full-Time • 7+ years exp • Wellington
Java
Python
SQL
Java
Spring Boot
Databases
Apache Kafka
Databricks
Neo4j
AI/ML
Flink
Spark
Frontend
GraphQL
DevOps
Azure
CI/CD
Datadog
Dynatrace
Kibana
Kubernetes
OpenShift
Platform Engineering
Splunk
Amazon ECS
Apply
$123k – $251k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Dallas • Denver • Birmingham
Java
SQL
Java
Gradle
Hibernate
Maven
Spring Boot
Spring Framework
Databases
Apache Kafka
MySQL
Redis
DevOps
CI/CD
Dynatrace
Jenkins
Kubernetes
OpenShift
Cybersecurity
SonarQube
Apply
$57k – $150k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • Paris
DevOps
Grafana
Kubernetes
Prometheus
Proxmox VE
Scaleway
HPC
Apply
$38k – $99k per year (Estimated) • Remote/Hybrid • 2+ years exp • Paris
DevOps
Scaleway
Management
Confluence
Jira
Slack
n8n
Apply
$86k – $184k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Paris
DevOps
Scaleway
HPC
Apply
$52k – $111k per year (Estimated) • Remote/Hybrid • Full-Time • Paris
DevOps
Scaleway
Apply
$12k – $14k per year • In office • Internship • Paris
Go
Databases
PostgreSQL
DevOps
Rest API
Scaleway
IAM
Apply
In office • Full-Time • PhD • Paris
Node JS
JavaScript
Node JS
Commander.js
Apply
$45k – $115k per year (Estimated) • In office • Full-Time • 3+ years exp • Paris
Apply
$35k – $87k per year (Estimated) • In office • Full-Time • 3+ years exp • Paris
Analytics
Power BI
Apply
$51k – $92k per year (Estimated) • Remote/Hybrid • Internship • 5+ years exp • Paris
Apply
$63k – $119k per year (Estimated) • Remote • Full-Time • 6+ years exp • PhD • Paris
Node JS
JavaScript
Node JS
Commander.js
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.