First seen by Alion on Sep 29, 2026.
The Observability Architect provides strategic leadership to define and advance enterprise-wide observability capabilities across the organisation. This role drives the adoption of modern observability practices to enhance system reliability, operational resilience, and performance visibility. It is responsible for establishing observability standards, governance frameworks, and scalable telemetry architectures while enabling engineering teams to proactively monitor, troubleshoot, and optimise complex distributed systems. The position also plays a key role in fostering observability excellence, improving service reliability metrics, and aligning technology initiatives with organisational objectives.
Responsibilities:
- Define and own the enterprise observability vision, roadmap, standards, and governance.
- Lead and mentor the Observability Engineering team.
- Design and implement enterprise-scale OpenTelemetry architectures.
- Establish standards for metrics, logs, traces, events, and business telemetry.
- Drive end-to-end distributed tracing and application performance monitoring.
- Define and govern SLIs, SLOs, error budgets, and alerting standards.
- Lead observability adoption across cloud, platform, infrastructure, and application teams.
- Optimise telemetry pipelines for scalability, reliability, and cost efficiency.
Requirements:
- Bachelor's degree in computer science or Information Technology.
- 12+ years in Software Engineering, Platform Engineering, SRE, or Enterprise Architecture.
- 5+ years leading observability or platform engineering teams.
- 3+ years of hands-on OpenTelemetry architecture experience.
- Strong expertise in distributed systems, cloud-native platforms, and enterprise monitoring.
Leadership Expectations:
- Build and lead a high-performing Observability Centre of Excellence.
- Collaborate across Engineering, SRE, Platform, Security, and Architecture teams.
- Drive enterprise-wide observability standards and governance.
- Improve operational resilience, MTTD, MTTR, and reliability metrics.
Technical Expertise:
- OpenTelemetry SDKs and Collectors.
- Distributed Tracing, Metrics, Logs, and Telemetry Pipelines.
- Kubernetes and Cloud Platforms (AWS, Azure, GCP).
- Grafana, Prometheus, Dynatrace, Datadog, Splunk, New Relic, Elastic, or AppDynamics.
- Terraform, GitOps, CI/CD, and DevOps Automation.
- Java,. NET, Python, Go, or Node.js .

