389,999open jobs
10,334companies
49,652added this week
Browse all
Salary
$87k – $207k per year (Estimated)
Location
In office (Singapore)
Seniority
Staff · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Mastercard is an American payments technology company whose origins date to 1966, when a group of banks formed the Interbank Card Association to compete with BankAmericard. Like its main rival it does not issue cards or extend credit; it operates the network that authorises, clears and settles transactions between issuing banks, acquirers and merchants in more than two hundred countries. Headquartered in Purchase, New York, the company has built a large services business alongside the core network, covering fraud and identity products through its Ethoca and RiskRecon acquisitions, open banking, consulting and loyalty programmes.

Our Purpose

Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.

Title and Summary

Lead, Site Reliability EngineerRole Overview

The Real Time Payments International team is looking for a Lead, Site Reliability Engineer to drive day-to-day operational stability and support the reliability of critical payment platforms by implementing automation, leverage best practices and work with a high-impact team responsible for driving production readiness, reliability, and DevOps automation across Mastercard platforms.

This role plays a key part in incident management, change readiness, and platform operations, while contributing to continuous improvement initiatives.

The ideal candidate brings strong technical troubleshooting skills, operational discipline, and ownership mindset, with the ability to lead during incidents and collaborate effectively across engineering and program teams.

Key Responsibilities

Platform Operations & Stability

  • Support end-to-end availability, monitoring, and performance of critical payment platforms.
  • Execute operational processes to ensure platform health and stability.
  • Participate in capacity checks, readiness validations, and environment monitoring.

Incident Management & Execution

  • Actively manage and coordinate incident triage and resolution.
  • Serve as incident commander driving medium to high-severity incidents.
  • Ensure timely updates, accurate impact assessment, and appropriate escalation.
  • Contribute to root cause analysis with clear identification of actions and ownership.

Change & Release Support

  • Participate in highlighting gaps and defining test cases required for a change in lower environments and validate lower environment test completeness.
  • Ensure adherence to change governance processes (test case reviews, checklists, approvals, rollback readiness).
  • Engage in creating change plans and support execution of production changes, deployments, and validations.

Technical Troubleshooting

  • Perform hands-on troubleshooting across:

o Application behaviour and dependencies.

o Infrastructure components (compute, network, storage).

o Database and performance issues.

  • Collaborate with engineering, infrastructure and other technical teams to isolate and resolve issues efficiently.

Monitoring & Observability

  • Improve system health monitoring using observability tools and alerts.
  • Identify gaps in alerting and contribute to improving quality of alerting and dashboards.
  • Ensure proactive detection of anomalies using observability tools.

Automation & Process Improvement

  • Contribute to automation initiatives to reduce toil and errors.
  • Identify repetitive operational tasks and drive improvements.
  • Support implementation of DevOps best practices.
  • Leverage AI-driven tools to improve monitoring, incident detection, and operational efficiency, enabling faster troubleshooting and reduced manual effort in day-to-day operations.

Stakeholder Coordination

  • Work closely with engineering, program teams, and external partners during incidents and changes.
  • Provide structured updates to stakeholders with clarity and consistency.
  • Ensure alignment during critical activities.

Risk Identification

  • Highlight operational and platform risks including test coverage gaps, infrastructure constraints, dependency risks.
  • Escalate issues proactively and support mitigation tracking.

Team Contribution & Mentorship

  • Support onboarding and guidance of junior team members.
  • Contribute to runbooks, documentation, and knowledge sharing.
  • Drive consistency in execution and adherence to operational standards.

Success in This Role Looks Like:

1. Deep Operational Ownership (“Built to Run” Mindset)

A successful Lead SRE Engineer is fully accountable for the operational health of their program, not just responsive to incidents.

  • Monitoring, alerting, and dashboards that reflect real customer impact.
  • Emergency response and incident leadership, including clear communications and post-incident follow-ups.
  • Capacity planning and readiness aligned with product and business growth.
  • Change management discipline, ensuring safe, compliant releases.

2.Strong Technical & System-Level Understanding

A Lead SRE Engineer is expected to operate at system dependency level, not just ticket or tool level.

  • Have a strong understanding of application business logic and workflows.
  • Have a clear grasp of upstream/downstream dependencies.
  • Expertise in observability (alerts, dashboards, synthetic monitoring).
  • Ability to drive automation to reduce manual toil and recurring issues.
  • End to End ownership of tasks and activities.

3. Incident Leadership & Decision-Making Under Pressure

Beyond technical skill, Leads are distinguished by how they lead during high-severity situations.

  • Takes command of major incidents, not waiting to be asked.
  • Maintains calm, structured communication with engineering, product, and leadership.
  • Balances speed vs risk in decision-making.
  • Ensures clear ownership of actions, timelines, and follow-ups.
  • Drives root cause analysis and systemic fixes, not just recovery.

4. Proactive Risk & Reliability Engineering

A successful Lead SRE prevents incidents more than fight them.

  • Identifies systemic risks before they become outages.
  • Pushes for design, monitoring, or process improvements.
  • Challenges “tribal knowledge” by insisting on documentation and runbooks.
  • Drives improvements aligned with operational maturity models.

5. Leadership Without Formal Authority

Lead SRE Engineers often lead without being people managers, which requires strong influence skills.

  • Mentors and coaches senior and mid-level SRE’s.
  • Sets the technical and behavioral bar for the team.
  • Gives clear, constructive feedback.
  • Acts as a role model for ownership, urgency, and professionalism.
  • Builds trust with Engineering, Product, and Platform teams.
  • Flexible in terms of working hours where needed.

6. Excellent Cross-Functional Communication

SRE Leads sit at the intersection of technology, operations, and business.

  • Translating technical issues into business impact to communicate clearly with senior stakeholders during incidents.
  • Setting expectations early and transparently with junior team members.
  • Represents SRE confidently in planning, reviews, and retrospectives,
  • Ensuring post-incident learnings are shared and acted upon.

All About You:

Bachelor’s degree in Computer Science, Engineering, or related field.

8+ years experience in production support, SRE, or BizOps roles.

Exposure to managing incidents and supporting distributed systems.

Experience in payments ecosystem will be preferred.

Knowledge of monitoring and alerting tools like Splunk , Dynatrace , Blaze meter.

Knowledge of automation and DevOps practices.

Experience working in cross-functional and high-pressure environments.

Ability to organize, multi-task and prioritise work based on current business needs.

Possesses strong verbal and written communication skills.

Strong relationship skills, collaborative skills and stakeholder management skills.

Experience in one or more of the following is preferred: C, C++, Java, Python, Go, Perl or Ruby.

Interest in designing, analysing and troubleshooting large-scale distributed systems.

Ability to work with little or no supervision.

Corporate Security Responsibility

All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:

  • Abide by Mastercard’s security policies and practices;

  • Ensure the confidentiality and integrity of the information being accessed;

  • Report any suspected information security violation or breach, and

  • Complete all periodic mandatory security trainings in accordance with Mastercard’s guidelines.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
389,999 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Singapore
SAP CPI Consultant 8 hours ago
$37k – $80k per year (Estimated) • In office • Full-Time • Warsaw
Java
DevOps
AWS
Rest API
VMWare
Management
ServiceNow
Marketing
Salesforce
Apply
$52k – $98k per year (Estimated) • In office • Full-Time • Lyon
JavaScript
TypeScript
Java
Java
Spring Boot
Spring Security
Databases
MySQL
PostgreSQL
Frontend
Vue.js
DevOps
Docker
Git
GitLab
Jenkins
Kubernetes
Rest API
Apply
$74k – $158k per year (Estimated) • In office • Full-Time • Barcelona
AI/ML
AI Agents
EU AI Act
Feature Store
Fine-tuning
Hallucination
Human-in-the-Loop
LLM
LLM Guardrails
RAG
DevOps
CI/CD
Incident Management
Apply
$50k – $105k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Basingstoke
Python
DevOps
Ansible
AWS
Incident Management
Management
ServiceNow
Apply
$21k per year • Remote • Moscow
Go
Java
Node JS
PHP
Python
SQL
JavaScript
Java
Hibernate
Spring Boot
Spring Cloud
Spring Data JPA
Spring MVC
Spring Security
Apply
$79k – $155k per year (Estimated) • In office • Full-Time • Dublin
Apply
$111k – $238k per year (Estimated) • In office • Full-Time • Dublin
Cybersecurity
NIST CSF
Apply
Software Engineer I 10 hours ago
$14k – $39k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Pune
DevOps
Incident Management
Apply
$21k – $54k per year (Estimated) • In office • Full-Time • Pune
PowerShell
Python
Node JS
JavaScript
Node JS
Prisma
DevOps
AIOps
Amazon CloudWatch
Amazon EC2
Amazon EKS
Amazon S3
ArgoCD
AWS
AWS Lambda
Azure
Azure AKS
Azure DevOps
Bicep
Chaos Engineering
CI/CD
CloudFormation
Datadog
Docker
Dynatrace
GitHub
GitHub Actions
GitLab
GitLab CI
GitOps
Grafana
Helm
IAM
Incident Management
Istio
Jenkins
JFrog Artifactory
Kubernetes
Kustomize
Linkerd
OpenShift
Progressive Delivery
Prometheus
Rest API
Self-Healing
Service Mesh
Splunk
Terraform
Cybersecurity
Prisma Cloud
Wiz
Apply
In office • Full-Time • Dubai
Apply
$38k – $84k per year (Estimated) • In office • Full-Time • 1+ year exp • Bachelor's Degree • Singapore
Apply
$150k – $250k per year • In office • Full-Time • 2+ years exp • Singapore
Python
AI/ML
AI Agents
LLM
Post-training
Reinforcement Learning
Synthetic Data
DevOps
Docker
Apply
$150k – $250k per year • In office • Full-Time • 2+ years exp • Singapore
Python
DevOps
Docker
Apply
$150k – $250k per year • In office • Full-Time • 2+ years exp • Singapore
Python
AI/ML
AI Agents
Reinforcement Learning
DevOps
Docker
Apply
$150k – $250k per year • In office • Full-Time • 5+ years exp • Singapore
Python
AI/ML
Post-training
Reinforcement Learning
Synthetic Data
DevOps
Docker
Apply
See all jobs
This is one of many
389,999 more open roles from verified company boards, updated every day.