368,530open jobs
9,432companies
50,439added this week
Browse all
Location
In office
Overview
Company
Impact
Profile match
Zensar Technologies provides digital engineering and information technology services. Its practices cover experience design, cloud, data and application modernisation. The company serves banking, manufacturing, retail and healthcare clients.

Job Description: Engineer Lead - Site Reliability Engineering (SRE)

Role Overview

The Engineer Lead - Site Reliability Engineering (SRE) will provide technical and thought leadership to ensure the reliability, resiliency, scalability, and observability of mission-critical platforms supporting Banking Solutions, Payments, and Capital Markets.

This role blends advanced SRE practices, resiliency and chaos engineering, and service health governance with people and technical leadership. The Engineer Lead will define reliability standards, mentor teams, and partner closely with Engineering, DevOps, Security, and Product stakeholders to embed reliability as a core product feature.

What You Will Be Doing

SRE Leadership & Strategy

Act as the technical lead and reliability champion, driving SRE best practices across multiple teams and platforms.

Define and evangelize reliability standards, principles, and operational excellence frameworks.

Guide teams in balancing feature velocity with system reliability using error budgets and SLO-driven decision-making.

Mentor and coach engineers in SRE, observability, incident management, and automation.

Service Health, SLI/SLO & Observability

Define, implement, and govern Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.

Establish standardized service health monitoring and reporting frameworks across platforms.

Design and maintain end-to-end observability solutions covering infrastructure, applications, APIs, and customer experience.

Drive reliability insights through dashboards, health scores, and executive-level metrics.

Resiliency Testing & Chaos Engineering

Lead resiliency engineering initiatives to validate system behavior under failure conditions.

Design and execute chaos engineering experiments to proactively identify weaknesses in architecture and operations.

Integrate resiliency testing into CI/CD pipelines and pre-production environments.

Partner with development and architecture teams to ensure systems are fault-tolerant, self-healing, and resilient by design.

Incident Management & Operational Excellence

Lead high-severity incident response efforts, providing clear technical and operational direction.

Establish and continuously improve incident management, escalation, and communication practices.

Drive blameless post-incident reviews, ensuring root causes are addressed and preventive actions are implemented.

Measure and improve operational KPIs such as MTTR, MTTD, and incident recurrence rates.

Automation, Platform Reliability & Cloud Operations

Champion automation-first approaches to reduce toil and manual intervention.

Oversee deployment pipelines, configuration management, and release reliability practices.

Guide teams on Infrastructure as Code (IaC), environment consistency, and cloud governance.

Ensure disaster recovery, backup, and failover strategies are tested and production-ready.

Cross-Functional Collaboration & Governance

Collaborate with Engineering, QA, DevOps, Security, Architecture, and Product teams to embed reliability into the SDLC.

Ensure platforms comply with security, regulatory, and audit requirements, especially in financial services environments.

Influence technical roadmaps to prioritize resiliency, stability, and customer experience.

Required Skills & Experience

Core Requirements

Strong experience with Core SRE practices, including system reliability, incident management, automation, and observability.

Hands-on expertise in resiliency testing and chaos engineering methodologies.

Proven experience designing and operating SLI / SLO / Error Budget frameworks at scale.

Deep understanding of distributed systems, microservices architectures, and cloud-native platforms.

Experience with cloud platforms (AWS, Azure, and/or Google Cloud).

Hands-on experience with Docker and Kubernetes.

Expertise in monitoring, observability, and logging tools, such as:

Prometheus, Grafana, Datadog

Splunk, ELK Stack

Strong background in incident management, post-mortem facilitation, and production support.

Proficiency in automation and scripting (Python, Bash, Terraform, Ansible).

Experience managing and improving CI/CD pipelines (Jenkins, GitLab CI/CD, Azure DevOps).

Ability to lead technical discussions, influence decisions, and communicate effectively with senior stakeholders.

Strong ownership mindset with accountability for service reliability and customer outcomes.

Nice to Have (SRE+ Skills)

Experience with Harness Chaos Engineering (CE) or similar chaos engineering platforms.

Programming experience in Java, particularly for debugging, performance analysis, or building internal SRE tooling.

Experience implementing self-healing and auto-remediation workflows.

Exposure to banking, payments, or capital markets domains.

Familiarity with chaos engineering maturity models and reliability governance practices.

What You Will Be Doing

SRE Leadership & Strategy

Act as the technical lead and reliability champion, driving SRE best practices across multiple teams and platforms.

Define and evangelize reliability standards, principles, and operational excellence frameworks.

Guide teams in balancing feature velocity with system reliability using error budgets and SLO-driven decision-making.

Mentor and coach engineers in SRE, observability, incident management, and automation.

Service Health, SLI/SLO & Observability

Define, implement, and govern Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.

Establish standardized service health monitoring and reporting frameworks across platforms.

Design and maintain end-to-end observability solutions covering infrastructure, applications, APIs, and customer experience.

Drive reliability insights through dashboards, health scores, and executive-level metrics.

Resiliency Testing & Chaos Engineering

Lead resiliency engineering initiatives to validate system behavior under failure conditions.

Design and execute chaos engineering experiments to proactively identify weaknesses in architecture and operations.

Integrate resiliency testing into CI/CD pipelines and pre-production environments.

Partner with development and architecture teams to ensure systems are fault-tolerant, self-healing, and resilient by design.

Incident Management & Operational Excellence

Lead high-severity incident response efforts, providing clear technical and operational direction.

Establish and continuously improve incident management, escalation, and communication practices.

Drive blameless post-incident reviews, ensuring root causes are addressed and preventive actions are implemented.

Measure and improve operational KPIs such as MTTR, MTTD, and incident recurrence rates.

Automation, Platform Reliability & Cloud Operations

Champion automation-first approaches to reduce toil and manual intervention.

Oversee deployment pipelines, configuration management, and release reliability practices.

Guide teams on Infrastructure as Code (IaC), environment consistency, and cloud governance.

Ensure disaster recovery, backup, and failover strategies are tested and production-ready.

Cross-Functional Collaboration & Governance

Collaborate with Engineering, QA, DevOps, Security, Architecture, and Product teams to embed reliability into the SDLC.

Ensure platforms comply with security, regulatory, and audit requirements, especially in financial services environments.

Influence technical roadmaps to prioritize resiliency, stability, and customer experience.

Required Skills & Experience

Core Requirements

Strong experience with Core SRE practices, including system reliability, incident management, automation, and observability.

Hands-on expertise in resiliency testing and chaos engineering methodologies.

Proven experience designing and operating SLI / SLO / Error Budget frameworks at scale.

Deep understanding of distributed systems, microservices architectures, and cloud-native platforms.

Experience with cloud platforms (AWS, Azure, and/or Google Cloud).

Hands-on experience with Docker and Kubernetes.

Expertise in monitoring, observability, and logging tools, such as:

Prometheus, Grafana, Datadog

Splunk, ELK Stack

Strong background in incident management, post-mortem facilitation, and production support.

Proficiency in automation and scripting (Python, Bash, Terraform, Ansible).

Experience managing and improving CI/CD pipelines (Jenkins, GitLab CI/CD, Azure DevOps).

Ability to lead technical discussions, influence decisions, and communicate effectively with senior stakeholders.

Strong ownership mindset with accountability for service reliability and customer outcomes.

Nice to Have (SRE+ Skills)

Experience with Harness Chaos Engineering (CE) or similar chaos engineering platforms.

Programming experience in Java, particularly for debugging, performance analysis, or building internal SRE tooling.

Experience implementing self-healing and auto-remediation workflows.

Exposure to banking, payments, or capital markets domains.

Familiarity with chaos engineering maturity models and reliability governance practices.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$29k – $70k per year (Estimated) • Remote/Hybrid • Full-Time • 11+ years exp • Bachelor's Degree • Hyderabad
Python
SQL
Databases
Databricks
Snowflake
DevOps
AWS
Azure
CI/CD
Apply
Data Analyst 10 hours ago
$38k per year (gross) • Remote/Hybrid • Full-Time • Bachelor's Degree • Bucharest
Python
SQL
AI/ML
Hadoop
DevOps
Azure
Analytics
Power BI
Apply
$111k – $216k per year (Estimated) • Equity • Remote • Full-Time • 6+ years exp • Bachelor's Degree • United States
Python
Ruby
Databases
PostgreSQL
DevOps
AWS
Chef
CI/CD
Configuration Management
Datadog
Git
Jenkins
Nagios
Amazon CloudWatch
GitLab
Cybersecurity
Crowdstrike
FedRAMP
Analytics
Tableau
ETL/ELT
Apply
$33k – $78k per year (Estimated) • Equity • Remote • Full-Time • 8+ years exp • Bachelor's Degree • India
Apex
JavaScript
Python
TypeScript
Apex
Copado
Lightning Web Components
AI/ML
AutoGen
CrewAI
Fine-tuning
Hallucination
LangChain
LangGraph
LlamaIndex
LLM
RAG
Semantic Kernel
Semantic Search
Synthetic Data
Vertex AI
Agentforce
AWS Bedrock AgentCore
Semantic Search
AI Agents
Model Context Protocol
DevOps
AWS
CI/CD
GitHub Actions
Jenkins
Vector
GitHub
Cybersecurity
Crowdstrike
Management
Slack
Marketing
Salesforce
Apply
$18k – $51k per year (Estimated) • Remote • Moscow
Bash
Python
Databases
ClickHouse
PostgreSQL
AI/ML
Feature Store
Hadoop
LLM
DevOps
Ansible
CI/CD
Docker
GitLab
Kubernetes
Apply
In office • 8+ years exp
Databases
Apache Kafka
ElasticSearch
OpenSearch
DevOps
Amazon CloudWatch
Amazon EC2
Amazon EKS
AWS
IAM
Azure
Azure DevOps
CI/CD
CloudFormation
Docker
Grafana
Helm
Incident Management
Jenkins
Kibana
Kubernetes
Prometheus
Splunk
Terraform
Apply
$18k – $44k per year (Estimated) • In office • Pune
Management
Jira
Slack
Marketing
Salesforce
Apply
$30k – $59k per year (Estimated) • In office • Bachelor's Degree • Bengaluru
Java
Python
Scala
SQL
Databases
Apache Kafka
Google BigQuery
AI/ML
Airflow
dbt
Flink
Spark
DevOps
CI/CD
GCP
Git
Google Cloud Run
Kubernetes
Terraform
Cybersecurity
GDPR
Analytics
ETL/ELT
Apply
$22k – $56k per year (Estimated) • In office • 7+ years exp • Pune
Go
PowerShell
Python
TypeScript
AI/ML
Copilot
DevOps
IAM
Azure
Azure AKS
Azure DevOps
CI/CD
FinOps
GitHub
Kubernetes
Terraform
Apply
SRE Engineer 1 day ago
$25k – $62k per year (Estimated) • In office • Bachelor's Degree • Bengaluru
JavaScript
Python
Databases
Memcached
Redis
DevOps
Akamai
Amazon CloudWatch
Amazon EC2
Amazon EKS
Amazon S3
ArgoCD
AWS
AWS CDK
IAM
Chaos Engineering
CI/CD
CloudFormation
Datadog
Dynatrace
Error Budget
GitOps
Incident Management
Istio
Jenkins
Kubernetes
Nginx
OpenTelemetry
Service Mesh
SLI/SLO/SLA
Splunk
Terraform
Cybersecurity
PCI DSS
SOC 2
Management
Confluence
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.