368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$73k – $108k per year
Location
Remote/Hybrid (Delft, Netherlands)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
TOPdesk helps organizations to improve their services with user-friendly standardized software. Our starting point is the needs of the organization, the people who will be using the software and the customers they will support. Collaboration is es...

TOPdesk builds service management software used across education, healthcare, government, and manufacturing. We are 700+ colleagues in 8 offices worldwide. Founded over 30 years ago, we serve more than 10 million users worldwide and have been helping organisations deliver better services ever since.

We are an open, collaborative organisation with little hierarchy - people own their work end to end and are trusted to make the decisions that matter. We are reinventing ITSM and ESM for the agentic era, building AI agents our customers can trust, and this role is part of that.

About the role

Our Azure SaaS estate keeps service management running for thousands of organisations worldwide, under SLA-backed 24/7 availability. As a Senior Site Reliability Engineer, you own the reliability of that estate as an engineering problem - you set the SLOs, engineer out the toil behind them, and make the platform faster to change and cheaper to operate without trading away resilience.

You sit in the SaaS infrastructure function, working alongside cloud engineering and the product squads shipping to production. You bring our AI-native ways of working into reliability: agents and bounded automation with observability, approvals, containment, and rollback - self-healing systems, not runbooks worked by hand.

What this is not A ticket-driven, break-fix ops role kept away from the code. This is reliability as engineering - you own SLOs and error budgets, automate what you repeat, and design the platform to recover itself rather than reacting incident by incident.

The team

We are a group of social technicians who value transparency, open feedback, and a healthy work-life balance - and who treat reliability as a shared, measurable objective, not a firefight.

What you'll own

  • SLOs and error budgets. Define and own service-level objectives across the Azure (and potentially multi-cloud) estate, and use error budgets to steer the balance between shipping change and protecting reliability.
  • Toil elimination and self-healing automation. Identify toil, classify it, and engineer it out - feeding self-healing automation and your findings into the reliability roadmap. Stand up an agent-based support layer that owns recurring toil and continuously feeds improvements back into reliability.
  • Observability consolidation. Standardise metrics, alerting, and tracing across all datacenters, close coverage gaps on cloud workloads, and measurably reduce the alert-to-incident ratio from baseline.
  • Incident response and blameless postmortems. Lead incidents to resolution, run blameless postmortems, and turn every learning into a durable fix or an automation candidate.
  • Reliability of releases. Harden CI/CD and progressive delivery - canaries, safe rollouts, automated rollback - so change velocity and reliability rise together.
  • Capacity and performance. Model capacity, load-test critical paths, and keep the platform within its performance envelope as it scales across regions.
  • AI-native reliability. Bring agents and bounded automation - with observability, approvals, containment, and rollback - into detection, diagnosis, and remediation.
  • Runbooks that get used. Every alert links to a runbook; every runbook links to an automation candidate. You leave things more legible than you found them.
  • Capacity and cost forecasting. Own capacity and cost planning across the multi-cloud estate, model usage and growth trends, and forecast short and long term infrastructure needs so spend and scaling decisions stay ahead of demand rather than reacting to it.

How you approach the work

  • Automate what you repeat - if you have done it manually twice, the third time is a design problem.
  • Measure before optimising: SLOs, baselines, and dashboards before opinions.
  • Design for failure - assume things break, and make recovery automatic and observable.
  • Consultative, not gatekeeping: you pair with product engineering teams and transfer knowledge as you go.
  • Treat cost and reliability as joint objectives, not a forced trade-off.
  • Pro-active collaboration with product teams. You are involved in the early phases of product development, including design to help the teams make optimal choices and timely introduce appropriate SRE practices.

Technical environment

  • Scale: 10+ global datacenters; SLA-backed, 24/7 multi-tenant SaaS serving millions of end users.
  • Cloud: Azure across all production regions, with a mature landing-zone and networking architecture.
  • Compute: Kubernetes / Azure AKS alongside traditional VM infrastructure, all managed as code.
  • Infrastructure as code: Terraform via CI/CD and GitOps workflows; configuration management with Puppet and Ansible across Linux and Windows.
  • Observability: metrics, alerting, and tracing across cloud-native and self-managed layers (e.g. Grafana, Prometheus, VictoriaMetrics, Influx).
  • Automation: Python and automation tooling - and we expect you to take the reliability stack to the next level, not just operate today's.
  • Legacy: Java, MS SQL, heritage architecture - being decomposed. The SRE role is not responsible for the Java application code.
  • How we build: Claude Code as our primary AI-native SDLC tool; subagents and multi-agent workflows; MCP tool integrations; shared prompt, agent, and eval libraries.

Success in your first year

  • SLOs and error budgets are defined for the estate's critical services and actively used to steer delivery decisions.
  • The alert-to-incident ratio is measurably down, and runbooks you wrote are used by the on-call shift without escalation.
  • Toil you identified is automated - or has a credible, documented roadmap to be - and self-healing covers at least one high-frequency failure mode.
  • Postmortems produce durable fixes, not repeat incidents; recurring-incident rate is trending down against a documented baseline.
  • Product squads consult you during design, not only after incidents.

Required

  • Proven hands-on experience (5+ years) as a Site Reliability, DevOps, or Infrastructure Engineer running a production cloud environment at scale (Azure).
  • Fluent with SLOs, error budgets, and reliability engineering practice - you have set them, not just read about them.
  • Strong observability skills at scale - Grafana, Prometheus, VictoriaMetrics, or equivalent - including alerting and tracing.
  • Kubernetes at operator level: Helm, namespace management, ingress controllers, RBAC, persistent volumes.
  • Coding for automation (Python or equivalent) and Terraform delivered via CI/CD.
  • Linux system administration - you understand what Puppet or Ansible is doing, not just whether it ran green.
  • Comfortable leading incidents in an on-call rotation with real SLA obligations, and the maturity to know when to escalate.
  • Strong written communication - your postmortems, runbooks, and architecture notes are unambiguous.

Nice to have

  • Experience with progressive delivery - canaries, feature flags, automated rollback.
  • Experience working within or migrating toward an Azure Cloud Adoption Framework or enterprise landing-zone structure.
  • Current, personal practice of AI-native software delivery (Claude Code or equivalent).
  • Experience with EU data residency / sovereign cloud requirements.

What's in it for you

  • A strong focus on personal development - including, in the Netherlands, our “10 to Grow” programme: 10% of your time and budget for your own growth.

  • A hybrid working environment built on freedom, trust, and responsibility.

  • An open, informal, and supportive culture, with collaboration across national borders.

  • Excellent employment conditions.

Want to apply?

Does this sound like your next step? Apply with your CV and a short motivation letter via the application form. Tell us about a service you made more reliable - the SLOs you set, the toil you engineered out, and the outcome (uptime, incidents reduced, or recovery time improved).

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Delft
$68k – $85k per year • In office • Full-Time • Master's Degree • San Jose
Python
AI/ML
AI Agents
DevOps
Amazon EC2
AWS
AWS Lambda
Bitbucket
CI/CD
CloudFormation
Docker
Git
Kubernetes
Terraform
Amazon S3
IAM
HPC
Cybersecurity
Least Privilege
Apply
$20k – $47k per year (Estimated) • In office • 3+ years exp • Almaty
Bash
Python
DevOps
Ansible
GitOps
KVM
OpenStack
OpenTofu
Terraform
VMWare
Amazon S3
GitLab
Apply
Cloud Engineer (AWS) 6 hours ago
In office • Full-Time • 5+ years exp • Bachelor's Degree • Dalian
Python
DevOps
AWS
CI/CD
CloudFormation
Docker
FinOps
GitHub Actions
Kubernetes
Terraform
GitHub
Apply
Mobile Engineer 1 day ago
$18k – $59k per year (Estimated) • In office • 7+ years exp • Mumbai
Java
Kotlin
Objective-C
Swift
Frontend
GraphQL
DevOps
CI/CD
Git
Management
Confluence
Jira
Apply
$19k – $47k per year (Estimated) • Remote • Full-Time • Perm
C#
C#
ASP.NET Core
Dapper
Entity Framework Core
Databases
Apache Kafka
ClickHouse
ElasticSearch
PostgreSQL
RabbitMQ
Redis
DevOps
CI/CD
Docker
Docker Compose
GitHub Actions
Kubernetes
TeamCity
GitHub
GitLab
Apply
$73k – $91k per year • Remote/Hybrid • Full-Time • 5+ years exp • Kaiserslautern
Python
SQL
TypeScript
Databases
PostgreSQL
AI/ML
AI Agents
Claude
Claude Code
Model Context Protocol
DevOps
Azure
CI/CD
Kubernetes
Terraform
Apply
Technical Support 12 days ago
$19k – $50k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • São Paulo
SQL
Apply
$149k – $156k per year • Remote/Hybrid • Full-Time • 5+ years exp • London
Apply
$142k – $156k per year • Remote/Hybrid • Full-Time • 5+ years exp • Manchester
Apply
Engineering Manager 12 days ago
$98k – $115k per year • Remote/Hybrid • Full-Time • Delft
Java
Python
SQL
TypeScript
Databases
MS SQL
PostgreSQL
AI/ML
AI Agents
Claude
Claude Code
LLM Guardrails
Model Context Protocol
DevOps
Azure
Kubernetes
Terraform
Apply
$42k – $96k per year (Estimated) • In office • Internship • 2+ years exp • Delft
Apply
Engineering Manager 12 days ago
$98k – $115k per year • Remote/Hybrid • Full-Time • Delft
Java
Python
SQL
TypeScript
Databases
MS SQL
PostgreSQL
AI/ML
AI Agents
Claude
Claude Code
LLM Guardrails
Model Context Protocol
DevOps
Azure
Kubernetes
Terraform
Apply
$87k – $105k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Delft
Apply
$76k – $181k per year (Estimated) • In office • Full-Time • 8+ years exp • Master's Degree • Delft
Python
DevOps
SLURM
HPC
Robotics
Digital Twin
Apply
$93k – $199k per year (Estimated) • In office • Full-Time • 10+ years exp • Master's Degree • Delft
DevOps
HPC
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.