1,433,388open jobs
83,596companies
217,057added this week
Browse all
Salary
$106k – $126k per year
Location
Remote (United Kingdom)
Seniority
Senior · 6+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 8, 2026. First seen by Alion on Sep 25, 2026.

Overview
Company
Impact
Profile match
IAM Cloud is an international software company that builds desktop and cloud software to make cloud IT better. Makers of Cloud Drive Mapper for Microsoft 365.

Introduction

IAM Cloud builds software that simplifies IT management in the cloud. We are a fully employee owned, bootstrapped company - which means we answer to our customers and our team, not investors. With around 35 employees, we serve more than 3,000 organisations worldwide.

Being employee owned shapes how we work. Decisions are long-term, the culture is collaborative, and the people who build the company share in its success. We have leaned deliberately into AI and agentic workflows across the business as a core part of how a focused team competes with companies ten times our size. We are fully remote with employees in the UK, Ireland, Spain, Germany and Argentina and have been remote for years, as a deliberate choice about how good work gets done.

About the role

IAM Cloud is an Azure-first SaaS company with a small, fully remote engineering team. We run our own platform end to end: infrastructure, pipelines, releases and production operations all sit with the same handful of engineers. That means ownership is broad, documentation matters, and reducing toil is treated as a first-class goal rather than a side project.

We recently added a Senior Cloud Platform Engineer to own our CI/CD and infrastructure-as-code foundations. This role is the reliability counterpart: the person who owns how we see, measure and respond to what production is doing. You will own monitoring, observability, alerting and on-call at IAM Cloud, as the person who sets the standards, writes the policy and makes the engineering team better at running its own services.

Today our telemetry lives in Azure Monitor, Log Analytics and Application Insights, with Azure Managed Grafana for dashboards and alerting. It works, but it has grown organically. Alert thresholds are inherited rather than designed, retention and cost are not governed, instrumentation is inconsistent across services, and our on-call tooling is reaching end of life and must be replaced on a fixed deadline in the first half of 2027. We want someone who can come in, take a clear-eyed look at all of that, and turn it into something deliberate.

Where we are heading

Over the medium term we intend to move more of our workloads onto Kubernetes (AKS), with Prometheus alongside Azure Monitor and OpenTelemetry as the instrumentation standard. That is a direction rather than a dated plan, and it is not the focus of this role in its first year. What matters now is that whoever owns observability here has a view on what good looks like on that platform, so that the standards you set today carry over cleanly when we get there.

What you will own (outcomes)

  • Define and evolve the observability strategy, target architecture and standards for the platform - what we instrument, how, where it lands, how long we keep it and what it costs.
  • Make Grafana the place engineers goto understand production: dashboards and alerting designed around services and customer journeys, not around whichever metrics happened to be available.
  • Establish service level objectivesfor our customer-facing services, with error budgets that engineering and product actually use to make decisions.
  • Rationalisealertingso that every page is actionable, owned and has a runbook - and so that the on-call engineer's phone is quiet when nothing is wrong.
  • Lead and champion the move to a modern incident response and on-call platform(incident.io or similar) integrated with Microsoft Teams, and design the on-call model around it: rota, escalation, and the post-incident review process that feeds back into the backlog.
  • Champion observability-first thinkingacross the engineering team, so developers instrument their own services well without needing you in the loop.

First-quarter deliverables

  • A written telemetry governance policycovering Log Analytics table plans, retention tiers, sampling and cardinality controls - with a baseline of current ingestion and spend and a target.
  • An alert rationalisationpass across Azure Monitor and Grafana: every remaining alert mapped to a service, an owner, a severity and a runbook; noisy or unowned alerts removed or downgraded.
  • Selection, delivery and cutover of a new incident response and on-call platform(incident.io or similar, Teams-integrated) to replace our current tooling before its end-of-life date, including parallel running and a documented on-call model.
  • SLOs and SLIsdefined for the top customer-facing services, with dashboards and burn-rate alerting.
  • Looking further out:as our Kubernetes and Prometheus direction firms up, you will own the observability standards for it - collector topology, log routing, dashboards and alerting - so that it lands well-instrumented from day one.

Day to day

  • Design and implement OpenTelemetryinstrumentation across our .NET services, and the collector pipelines that shape, sample and route telemetry.
  • Own Azure Monitor, Log Analytics and Application Insights configuration, managed as code alongside the rest of our estate (Bicep first).
  • Own Grafana end to end: data sources, dashboards-as-code, alert rules and the conventions the team follows when adding to them.
  • Run the incident response and on-call platform once it is in place: routing, escalation policies, integrations with Azure Monitor and Grafana alerting, and the Teams workflow around an incident.
  • Contribute observability requirements and standards to our Kubernetes and Prometheus direction as it takes shape, alongside the platform engineer who owns that work.
  • Build and maintain Grafana dashboards and alert rules that answer real operational questions, and retire the ones that don't.
  • Lead incident response when it matters, run blameless post-incident reviews, and turn findings into engineering work.
  • Track and reduce reliability metrics that matter: MTTR, alert volume per on-call shift, incident recurrence, telemetry cost per service.
  • Evaluate tooling and AI-assisted operations capabilities on their merits, and give the team a clear recommendation rather than a shortlist.
  • Write things down: standards, runbooks, decision records, and the "why" behind thresholds.
  • Take part in a light-touch on-call rotation with the rest of the engineering team.

What we're looking for

What we're looking for (must-haves)

  • 6+ years in engineering, with 3+ in a reliability, observability or production-operations role where you owned outcomes rather than executed tickets.
  • You have defined SLOs, SLIs and error budgets for real services and can talk through how they changed decisions.
  • You have reduced alert noise - you can describe an alerting estate you inherited, what you cut, what you kept, and how you knew it was safe.
  • Hands-on OpenTelemetry:instrumenting services, running collectors, and making deliberate choices about sampling and cardinality.
  • Telemetry cost and retention governance: you understand that observability has a bill, and you have managed it.
  • Deep Azure Monitor / Log Analytics / Application Insights experience, including KQL, and strong Grafana skills - dashboards, alerting and managing both as code.
  • Infrastructure as code - Bicep is our standard; strong Terraform experience is fully transferable.
  • Clear written and spoken communication at C1 level or above, or native-level business English. You will be setting policy for a remote team; if it isn't written down, it doesn't exist.

Nice to have

  • Experience selecting, implementing or migrating incident-management and on-call platforms - strongly desirable given the first-quarter deliverables.
  • Prometheus-based monitoring and alerting, including exporters, recording rules and cardinality management.
  • Kubernetes in production, ideally AKS, and the observability patterns that go with it.
  • A software development background, ideally .NET, so instrumentation conversations with developers are peer to peer.
  • Familiarity with AI-assisted operations and coding tools, and a considered view on where they help and where they don't.
  • Azure certifications (AZ-104, AZ-400, AZ-305) or the CKA.
  • Experience in a small or scale-up environment where you were the observability function.

Why us?

How we work

  • Fully remote across the UK. We meet in person a few times a year for team events, but day-to-day work is remote.
  • Small, senior team. You will work directly with engineering leadership and have real input into technical decisions.
  • Async-friendly. We use writing as our default mode of communication and avoid meetings when a written update will do.
  • On-call is shared and light. We invest in making the platform boring rather than relying on heroics.

Our Offer

£80,000 - £95,000 salary (dependent on skills & experience)

Guaranteed £2,000 pay rise every year you're with us, separate from any merit or promotion increases.

Time off and flexibility

  • 26 days holiday plus public holidays, with the option to swap public holidays for days that are more meaningful to you.
  • Your birthday and work anniversary day off, every year.
  • Up to 4 weeks per year working from anywhere in the world.
  • Genuinely flexible, fully remote working - we trust you to manage your own time.

Health, family, and wellbeing - UK employees

  • Generous Becoming a Parent leave, designed to support every kind of family.
  • Medical scheme including remote GP, dental, optical, and diagnostics.
  • Access to Support Room - confidential counselling, therapy, and coaching.
  • Help@Handemployee assistance programme.
  • 5x salary life insurance through Unum.

Money and growth

  • Up to 10% pension match (UK employees).
  • Annual learning and development budget per function.

Interview process

We respect your time and aim to keep this efficient. Our typical process looks like this:
  • Initial conversation (45-60 minutes): a chat with our Head of People about you, the role, and any questions you have.
  • Technical conversation (60-90 minutes): a discussion with a couple of our senior engineers about your experience, our stack, and how you've approached real problems. This will also include a short, scoped live exercise relevant to the work you'd actually do here.
  • Final conversation (45-60 minutes): meet our CIO and one or two future team-mates to talk about the team, the company, and the road ahead.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,433,388 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
In your city
$63k per year • Remote (likely DACH) • Full-Time • Wels
DevOps
VMWare
Apply
≈ $81k – $186k per year (Estimated) • Remote (United States) • Contractor
SQL
Databases
Snowflake
Databricks
MS SQL
DevOps
Azure
Windows
Analytics
SSIS
SSAS
Apply
Azure Architect 2 hours ago
≈ $114k – $239k per year (Estimated) • Remote (United States) • Full-Time
DevOps
Azure
Apply
≈ $88k – $201k per year (Estimated) • Remote (United States) • TS/SCI • 5+ years exp • Bachelor's Degree
Python
Databases
PostgreSQL
Amazon Aurora
DevOps
Terraform
GitHub Actions
Terragrunt
CircleCI
Datadog
Azure
CI/CD
AWS
Docker
Kubernetes
Grafana
Configuration Management
Atlantis
Amazon ECS
Apply
≈ $126k – $240k per year (Estimated) • Remote (United States) • Full-Time • 4+ years exp
Python
JavaScript
TypeScript
Node JS
Node JS
Commander.js
Databases
PostgreSQL
ClickHouse
AI/ML
Quantization
Function Calling
LLM
OpenRouter
Prompt Caching
Tool Use
DevOps
GCP
Vercel
Cloudflare
SLI/SLO/SLA
API Gateway
Apply
≈ $57k – $141k per year (Estimated) • In office • Birmingham
PowerShell
C#
Databases
Databricks
Azure SQL Database
DevOps
Splunk
Terraform
Azure DevOps
Azure
CI/CD
Docker
Kubernetes
DNS
Cybersecurity
GDPR
Management
Agile
Scrum
Apply
≈ $80k – $156k per year (Estimated) • Hybrid • 5+ years exp • United Kingdom
SQL
C#
C#
.NET
DevOps
Azure DevOps
Azure
Apply
≈ $43k – $82k per year (Estimated) • Remote (Canada) • 3+ years exp • Ottawa
SQL
PowerShell
C#
C++
DevOps
VMWare
Azure
Windows Server
AWS
Hyper-V
Windows
Management
Intercom
Apply
≈ $37k – $77k per year (Estimated) • Remote (United Kingdom, Ireland) • United Kingdom
SQL
C#
C#
.NET
Management
Jira
ITIL
Apply
≈ $70k – $140k per year (Estimated) • In office • PhD • Blackburn
DevOps
VMWare
Azure
Proxmox VE
Hyper-V
Windows
VPN
MPLS
Cybersecurity
Okta
Crowdstrike
Apply
See all jobs
This is one of many
1,433,388 more open roles from verified company boards, updated every day.