368,634open jobs
9,437companies
50,578added this week
Browse all
Location
In office (Kuala Lumpur)
Seniority
Architect · 10+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
AIA Group is a prominent multinational insurance and finance organization headquartered in Hong Kong, serving markets across the Asia-Pacific region. The company provides a wide array of services, including life, accident, and health insurance, as well as savings and wealth management solutions. With a focus on promoting holistic well-being, it is dedicated to helping individuals and families across its diverse markets live healthier, longer, and better lives.

Are you ready to shape a better tomorrow?

AIA Digital+ is a Technology, Digital and Analytics innovation hub dedicated to powering AIA to be more efficient, connected and innovative as it fulfils its Purpose to help millions of people across Asia-Pacific live Healthier, Longer, Better Lives.

If you are hungry and driven to play an active role in shaping a better tomorrow, we want to hear from you. Because the work we do at AIA Digital+ makes a difference in the lives of millions of people, every day. We will equip you with the critical skills, tools and technology, and endless opportunities to learn, contribute and thrive in a dynamic and exciting environment.

If you want to shape a brighter future at AIA Digital+, please read on.

About the Role

The Observability Architect is responsible for defining, designing, implementing, and governing enterprise observability capabilities across AIA multi-cloud, hybrid, and containerized environments. The role will establish consistent monitoring, logging, tracing, event correlation, service health, and operational intelligence practices across Azure, Alibaba Cloud, and related enterprise platforms.

Dynatrace will be the primary observability platform, integrated with ServiceNow ITOM to support event management, incident enrichment, service mapping, root cause analysis, and operational automation. The role will also define complementary patterns for logging and open-source monitoring platforms such as Elastic, OpenSearch, Prometheus, Grafana, and OpenTelemetry.

A key objective of this role is to drive AIOps-enabled operations and self-healing automation to reduce alert noise, improve Mean Time To Detect (MTTD), reduce Mean Time To Resolve (MTTR), and improve overall platform and application reliability.

Roles and Responsibilities:

Observability Architecture & Design

  • Define and maintain enterprise observability reference architecture, standards, patterns, and governance across AIA Group and Business Units.
  • Design end-to-end observability for infrastructure monitoring, application performance monitoring, distributed tracing, logging, digital experience monitoring, network observability, and service health dashboards.
  • Establish observability KPIs, SLOs, SLIs, error budgets, alerting principles, and service reliability reporting standards.
  • Review solution designs and ensure observability requirements are embedded from architecture and delivery stages.

Dynatrace Platform Architecture

  • Lead architecture and governance of Dynatrace as the enterprise observability platform.
  • Design monitoring standards for applications, Kubernetes, virtual machines, databases, middleware, APIs, network services, and cloud-native services.
  • Define Dynatrace tagging standards, management zones, dashboard patterns, service mapping, synthetic monitoring, real user monitoring, and Davis AI adoption.
  • Drive platform configuration, onboarding patterns, operational dashboards, reporting, and observability data retention standards.

ServiceNow ITOM Integration

  • Architect integration between Dynatrace and ServiceNow ITOM Event Management / ITOM Health capabilities.
  • Design event correlation, alert enrichment, topology-aware service impact analysis, and automated incident creation workflows.
  • Define noise reduction, deduplication, priority mapping, escalation, and operational ownership models.
  • Support integration of observability insights into ITSM processes, command center dashboards, and operational reporting.

AIOps & Auto-Healing Automation

  • Define and implement AIOps operating patterns for anomaly detection, predictive alerting, root cause identification, and closed-loop remediation.
  • Design auto-healing workflows integrated with Dynatrace, ServiceNow ITOM, automation platforms, and cloud-native services.
  • Identify repeatable operational failure patterns and convert them into automated remediation runbooks where appropriate.
  • Drive continuous improvement to reduce manual intervention, alert fatigue, incident recurrence, MTTD, and MTTR.

Logging & Observability Data Platforms

  • Architect centralized logging and log analytics solutions using Elastic, OpenSearch, cloud-native log services, and enterprise logging standards.
  • Define log collection, normalization, enrichment, indexing, retention, access control, and cost governance standards.
  • Establish patterns for correlation across logs, metrics, traces, events, configuration, and service topology.
  • Ensure logging platforms support operational troubleshooting, security visibility, auditability, and compliance requirements.

Open Source Monitoring & Cloud Native Observability

  • Design monitoring and visualization solutions using Prometheus, Grafana, OpenTelemetry, and related open-source monitoring platforms.
  • Define observability patterns for AKS, Alibaba Cloud ACK, containers, microservices, APIs, service mesh, and DevOps pipelines.
  • Establish metrics federation, dashboarding, alerting, and integration approaches with Dynatrace and enterprise platforms.
  • Promote consistent instrumentation and telemetry standards across modern application architectures.

Azure & Alibaba Cloud Native Monitoring

  • Define monitoring standards for Azure platform services including Azure Monitor, Log Analytics, Application Insights, Network Watcher, Azure Managed Prometheus, Azure Managed Grafana, Azure Advisor, and Azure Resource Health.
  • Define monitoring standards for Alibaba Cloud services including CloudMonitor, Log Service (SLS), ActionTrail, ARMS, Managed Service for Prometheus, Security Center, and related observability and alarm services.
  • Ensure cloud-native monitoring capabilities are integrated with enterprise observability, ITOM, incident, and reporting processes.
  • Drive cost-aware telemetry design across Azure and Alibaba Cloud environments.

Security, Governance & Compliance

  • Embed security-by-design, least-privilege access, data protection, and compliance requirements into observability architecture.
  • Support audits, risk assessments, regulatory reviews, and evidence requirements related to monitoring, logging, and service reliability.
  • Define guardrails for observability platform access, retention, data classification, dashboard sharing, and operational reporting.
  • Maintain observability documentation, runbooks, standards, design patterns, and operational controls.

Stakeholder & Technical Leadership

  • Act as the enterprise subject matter expert for observability, monitoring, logging, AIOps, and self-healing automation.
  • Collaborate with Cloud Architecture, Cloud Engineering, Operations, Security, DevOps, Application, Service Management, and Business Unit teams.
  • Mentor engineering and operations teams on observability best practices, platform onboarding, dashboarding, and incident reduction techniques.
  • Drive observability transformation initiatives and promote SRE-aligned operational practices across the organization.

Minimum Job Requirements:

Skills:

  • 10+ years relevant experience in Cloud Architecture, Infrastructure, Operations, Observability, or Application Performance Management.
  • 5+ years practical experience designing and governing enterprise observability platforms.
  • Strong hands-on experience with Dynatrace, including APM, infrastructure monitoring, Kubernetes monitoring, dashboards, management zones, service mapping, and Davis AI capabilities.
  • Experience integrating observability platforms with ServiceNow ITOM / Event Management / ITSM processes.
  • Strong experience with logging platforms such as Elastic Stack, OpenSearch, or equivalent enterprise log analytics solutions.
  • Experience with open-source monitoring and visualization platforms such as Prometheus, Grafana, and OpenTelemetry.
  • Hands-on experience with Azure native monitoring services including Azure Monitor, Log Analytics, Application Insights, Network Watcher, Azure Managed Prometheus, and Azure Managed Grafana.
  • Hands-on experience with Alibaba Cloud native monitoring services including CloudMonitor, Log Service (SLS), ActionTrail, ARMS, Managed Service for Prometheus, and related observability services.
  • Strong understanding of Kubernetes, containers, microservices, APIs, network monitoring, distributed tracing, cloud networking, IAM, and security principles.
  • Experience implementing AIOps, event correlation, anomaly detection, automated remediation, and self-healing operations.
  • Experience working in enterprise or regulated environments is highly desirable.
  • Relevant professional certifications such as Dynatrace Associate/Professional, Microsoft Azure Solutions Architect Expert, Alibaba Cloud Professional Architect, CKA, ITIL, SRE Foundation, or TOGAF will be an advantage.
  • Sound understanding of IT partner ecosystem and partner collaboration in a multinational corporation.
  • Experience in top-tier multinational corporation will be an advantage.
  • Strong problem-solving and analytical skills.
  • Excellent communication, stakeholder management, and collaboration skills.
  • Ability to translate complex operational telemetry into actionable service reliability insights.
  • Ability to operationalize disruptive technology services, including building implementation roadmaps for observability, AIOps, and self-healing automation.

Build a career with us as we help our customers and the community live healthier, longer, better lives.

You must provide all requested information, including Personal Data, to be considered for this career opportunity. Failure to provide such information may influence the processing and outcome of your application. You are responsible for ensuring that the information you submit is accurate and up-to-date.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Kuala Lumpur
$181k – $367k per year (Estimated) • In office • Jersey City
DevOps
Azure
GCP
Cybersecurity
Least Privilege
Zero Trust
Apply
$169k – $321k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Phoenix
AI/ML
AI Agents
Anomaly Detection
LLM Guardrails
DevOps
AWS
Kong
Amazon S3
API Gateway
Cybersecurity
Zero Trust
Apply
$140k – $225k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree • Seattle
C++
Go
Python
Rust
Python
FastAPI
Databases
Neo4j
pgvector
Qdrant
PostgreSQL
AI/ML
LLM
SGLang
vLLM
Knowledge Graph
AI Agents
DevOps
AWS
Azure
Bicep
GCP
Karpenter
KEDA
Kubernetes
OpenTelemetry
OpenTofu
Terraform
Vector
Cybersecurity
Least Privilege
Apply
$191k – $267k per year • Equity • Remote • Full-Time • 4+ years exp • Master's Degree
Python
SQL
AI/ML
Anomaly Detection
Apply
Remote • Full-Time • 3+ years exp • Cairo
C#
JavaScript
TypeScript
C#
.NET
AI/ML
Copilot
DevOps
Azure
Azure DevOps
CI/CD
Rest API
Analytics
ETL/ELT
Management
Power Apps
Power Automate
Apply
In office • Full-Time • Kuala Lumpur
AI/ML
Copilot
DevOps
Azure
Azure DevOps
Design
Figma
Sketch
Management
Jira
UiPath
Apply
Graphic Designer 4 days ago
In office • Full-Time • Auckland
Design
Figma
Apply
In office • Full-Time • Melbourne • Sydney
Management
Confluence
Jira
Apply
Cloud Architect 4 days ago
$23k – $54k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Makati
PowerShell
Python
SQL
Databases
Azure Cosmos DB
Azure SQL Database
DevOps
Azure
Azure DevOps
Bicep
CI/CD
Docker
Jenkins
Kubernetes
Terraform
Apply
$69k – $160k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Singapore
Apply
Remote/Hybrid • Full-Time • 1+ year exp • Kuala Lumpur
Apply
In office • Bachelor's Degree • Kuala Lumpur
Apply
In office • Full-Time • 10+ years exp • Kuala Lumpur
COBOL
SQL
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Kuala Lumpur
ABAP
ABAP
ABAP Objects
DevOps
Rest API
SLI/SLO/SLA
Apply
In office • Full-Time • 8+ years exp • Bachelor's Degree • Kuala Lumpur
ABAP
DevOps
SLI/SLO/SLA
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.