Location
In office (Hong Kong)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
FWD Group is a pan-Asian life insurance provider headquartered in Hong Kong and established in 2013. The company offers a range of products including life and medical insurance, general insurance, employee benefits, and family takaful solutions. It operates across ten markets in Asia, serving millions of customers through a multi-channel distribution network that includes agents, bank partners, and digital platforms.
About FWD Group
FWD Group (1828.HK) is a pan-Asian life and health insurance business that serves approximately 40 million customers across 10 markets, including BRI Life in Indonesia. FWD’s customer-led and tech-enabled approach aims to deliver innovative propositions, easy-to-understand products and a simpler insurance experience. Established in 2013, the company operates in some of the fastest-growing insurance markets in the world with a vision of changing the way people feel about insurance. FWD Group is listed on the main board of the Hong Kong Stock Exchange under the stock code 1828.
For more information, please visit www.fwd.com
Purpose
- Own the Group-wide strategy, policy, and outcomes for IT resilience, service reliability, and platform modernization across FWD Group’s infrastructure, cloud platforms, and core system applications
- Set reliability objectives, govern error-budget policy, and hold decision rights over production change risk for NB and Customer facing related services
- Chair the Group Resilience Council and drive a federated SRE operating model across Group Office and Business Units (BUs)
- Strategize, Design, enforce and govern the group resilience standard, all systems must be Highly Availability, with DR plan, with failover plan
- Lead modernization of legacy platforms and production services by defining target-state architectures, resilience patterns, upgrade roadmaps, and remediation priorities to improve availability, scalability, security, and maintainability
- IT SRE provides advice to different teams on how to fix P1/P2 RCAs, drives troubleshooting / analyze / identify where in the code need to be fixed and oversee the entire troubleshooting process
- Drive modernization through observability, automation, SRE practices, and engineering enablement, ensuring incident learnings translate into platform hardening, architectural simplification, and faster, safer delivery
- Act as multi-SME / Generalist team, SMEs of different areas ( Security, Network, Cloud, Application, Infrastructure, etc)
- IT SRE manages and owns the new P1/2 escalation protocol
Key accountabilities
Modernization
- Define and drive the modernization roadmap for core system applications, including lifecycle management, upgrade strategy, technical debt reduction, platform simplification, and resilience-by-design requirements
- Lead modernization reviews for core systems to assess architecture fitness, recoverability, scalability, supportability, and security, and translate incident learnings into prioritized remediation and refactoring plans
- Establish modernization guardrails for core application estates covering observability, automation, release engineering, resilience patterns, decommission planning, and adoption of cloud-native or platform-standard capabilities where appropriate
- Enterprise governance & policy: Develop and own SRE Standards, Error Budget Policy, On-Call & Incident Command framework; enforce release gates based on reliability risk
- Enterprise governance & policy: Develop and own Group standards covering resilience, modernization guardrails,error budgets, on-call and incident command frameworks, and production change controls; enforce release gates based on reliability and modernization risk
- Reduce MTTR and incident recurrence; scale SLO coverage to ≥ 90% of critical services
- Drive reliability-by-design reviews for NB and Customer Facing system changes, preventing recurrence through architectural guardrails and automated release gates
- Platform ownership: Product-own Observability & AIOps platforms and drive modernization enablers including telemetry standards, engineering guardrails, automated release controls, and reusable patterns for cloud-native adoption
- Resilience: Approve DR tiers (RTO/RPO), lead chaos/DR exercises, and ensure cyber-resilience alignment with Security
- Financials: Own SRE platform budget, FinOps targets, and vendor SLA outcomes
- Org & talent: Build a global SRE leadership bench, run the SRE Academy, and operate a follow-the-sun model with healthy on-call
- Stakeholder engagement: Prepare and present executive reporting to GMT on production reliability risk and major incidents
- Provide leadership and decisioning across Group IT, local BUs IT, local BUs users and Group Digital & Data team by providing troubleshooting direction and approach
- Provide thought leadership in performing root case analysis and develop long-term prevention measures
- Governance Framework : Develop publish Group SRE Standards including SLO/SLI definitions, error budgets, release gates, on-call health, and post-incident review policies, track and monitor the standard is executed across Group IT, local BU, Group Digital & Data and the relevant stakeholders
- Engage and intimately involved in technical leadership decision-making and collaboration with other key Technology leaders within Group and local BUs in setting the governance framework in enhancing SRE and production reliability
- Embedded Chapters : Create embedded SRE chapters within each BU and function (e.g., Group Digital), supported by a central platform SRE team. This ensures local ownership while maintaining global consistency
- Accountable to maintain and comprehensive understanding of the multi facets and disciplines that underly the importance of maintaining SRE and production reliability of our platforms, systems and applications in order to achieve our business goal. This will require an underlying understanding of business principles, related to IT solutions, regulatory compliance, and data security compliance.
- Executive Dashboards and Escalation Protocols : Create executive dashboards and present to GMT and Excos on production risk and customer impact, collect feedback from GMT and Excos to strive for continuous improvement
- Design and responsible for a tiered escalation protocol for P1/P2 incidents (CTO → Senior MD → CEO) to ensure visibility and urgency
- Training and Culture: Operate the SRE Academy to build internal capability and promote a shared understanding of reliability principles
- Promote blameless post-incident reviews (PIRs) to foster trust and continuous improvement
Qualifications / Experience
- Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering, or related field
- 15+ years in large-scale infrastructure/platform/SRE, with 8-10+ years leading managers across multi-region operations
- Proven track record leading enterprise resilience and modernization transformations, including legacy remediation, cloud adoption, platform standardization, and operating model change in multinational environments
- Budget ownership and vendor/commercial leadership; experience reporting to GMT on production risk
- Deep expertise with multi-cloud (AWS/Azure/GCP), Kubernetes, CI/CD, Open Telemetry, and enterprise observability
- Familiarity with ITIL 4, ISO 22301/27001, incident response (NIST 800-61), and audit frameworks relevant to insurance
Knowledge & Technical skills
- SRE practices at scale (SLO/SLI design, error budgets, chaos engineering, capacity management)
- Observability stacks and AIOps (e.g., Elastic, Datadog, Dynatrace; correlation and anomaly detection)
- Strong understanding of cloud platforms, container orchestration, CI/CD, API and integration patterns, infrastructure-as-code, and modernization approaches such as refactoring, replatforming, and decomposing legacy services.
- Understanding of IT operations in Insurance, BCP/DR, and regulatory expectations
- Excellent communication and stakeholder management across Group and BUs; fluent in English
- Chinese written proficiency preferred
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,910 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Free forever. No card. Under a minute.
Your match
How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.
Recommended for you based on this role
Similar stack
Same company
Hong Kong
SDET
1 day ago
≈ $12k – $39k per year (Estimated) • In office • 4+ years exp • Gurgaon
Java
DevOps
AWS
Azure
CI/CD
Jenkins
QA
Appium
JMeter
Playwright
Rest-Assured
Selenium
Apply
Fullstack Developer (React/NodeJS)
1 day ago
≈ $18k – $46k per year (Estimated) • Remote • Full-Time • Tula
C#
JavaScript
Node JS
SQL
TypeScript
C#
.NET
Node JS
InversifyJS
Databases
DynamoDB
MySQL
AI/ML
Claude
Copilot
Cursor
OpenAI Codex
Frontend
Angular
React.js
Tailwind CSS
Mobile
Dependency Injection
DevOps
AWS
AWS Lambda
CI/CD
OpenTelemetry
Rest API
Terraform
Amazon CloudWatch
Amazon S3
API Gateway
GitHub
Cybersecurity
HIPAA
Apply
≈ $19k – $53k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Pune
Bash
JavaScript
Python
TypeScript
Frontend
Angular
React.js
DevOps
AWS
Azure
Datadog
Docker
GCP
Grafana
Kubernetes
Prometheus
Splunk
IAM
Cybersecurity
Keycloak
Apply
Senior Data Analyst
1 day ago
≈ $20k – $45k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Gurgaon
Python
SQL
Python
pySpark
Databases
Databricks
Microsoft Fabric
AI/ML
Hadoop
Spark
DevOps
AWS
Azure
Analytics
Power BI
Tableau
Apply
≈ $67k – $173k per year (Estimated) • In office • Contractor • 3+ years exp • Bachelor's Degree • Singapore
JavaScript
Python
DevOps
AWS
Azure
GCP
Cybersecurity
ISO 27001
OWASP Top 10
Apply
Manager, Data Analytics
11 days ago
In office • Full-Time • 5+ years exp • Ho Chi Minh City
Python
SQL
Analytics
Power BI
ETL/ELT
Apply
Apply
IT Business Analysis Senior Specialist
15 days ago
≈ $15k – $34k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Taguig
Management
Jira
Slack
QA
Appium
Selenium
Apply
Manager, Web Application & QA Assurance
15 days ago
≈ $65k – $162k per year (Estimated) • In office • Full-Time • 6+ years exp • Bachelor's Degree • Singapore
JavaScript
Node JS
SQL
TypeScript
Node JS
Nest.JS
AI/ML
Copilot
LLM
Ray
Frontend
GraphQL
Next.js
React.js
DevOps
Azure
Azure DevOps
CI/CD
GitLab CI
Jenkins
Kong
Rest API
Self-Healing
GitHub
GitLab
QA
Appium
Cypress
Gatling
JMeter
k6
Playwright
Postman
Rest-Assured
Selenium
SoapUI
Apply
Director, Data Scientist
15 days ago
≈ $94k – $227k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Singapore
Python
SQL
Databases
Databricks
AI/ML
AI Agents
Claude
Claude Code
CNN
Copilot
LLM
Prompt Engineering
Spark
Context Engineering
DevOps
GitHub
Apply
Application Support Manager, Asia Pacific
5 hours ago
In office • 5+ years exp • Bachelor's Degree • Hong Kong
Apply
Apply
Apply
AI Engineer | Insurance
1 day ago
In office • 3+ years exp • Hong Kong
Java
Node JS
Python
SQL
JavaScript
AI/ML
BERT
Fine-tuning
Hadoop
LangChain
LangGraph
LLM
RAG
Spark
Edge AI
Knowledge Graph
AI Agents
Google ADK
Frontend
React.js
DevOps
AWS
Azure
CI/CD
Docker
GCP
Git
Kubernetes
Cybersecurity
GDPR
Apply
AI Engineer | Financial Services
1 day ago
In office • 3+ years exp • Hong Kong
Java
Node JS
Python
SQL
JavaScript
AI/ML
AI Agents
BERT
Edge AI
Fine-tuning
Google ADK
Hadoop
Knowledge Graph
LangChain
LangGraph
LLM
RAG
Spark
Frontend
React.js
DevOps
AWS
Azure
CI/CD
Docker
GCP
Git
Kubernetes
Cybersecurity
GDPR
Apply
This is one of many
368,910 more open roles from verified company boards, updated every day.

