707,525open jobs
41,962companies
101,146added this week
Browse all
Salary
$140k – $233k per year
Location
Remote/Hybrid (Buffalo, United States)
Seniority
Staff · 10+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
M&T Bank is an American regional bank founded in Buffalo in 1856 to serve manufacturers and traders, and it has grown into one of the larger commercial banks in the northeastern United States. It concentrates on commercial and business banking, commercial real estate lending and retail branches across New York, the mid-Atlantic and New England, a footprint substantially enlarged by its 2022 acquisition of People's United Financial. Headquartered in Buffalo and listed on the New York Stock Exchange, it has a long-standing reputation for conservative credit underwriting through economic cycles.

Overview

The Site Reliability Engineering (SRE) Manager leads teams responsible for the reliability, availability, performance, and operational excellence of critical business applications and platforms. This role combines engineering leadership with deep expertise in production operations, observability, automation, incident management, and cloud technologies.

The SRE Manager partners with Engineering, Architecture, Infrastructure, Security, Product, and Business stakeholders to ensure systems are resilient, scalable, secure, and supportable. The role is accountable for driving operational excellence through automation, reliability engineering practices, and continuous improvement while developing high-performing SRE and Production Support teams.

Primary Responsibilities

Reliability & Operational Excellence

  • Define and execute SRE strategies that improve system reliability, availability, scalability, and performance.
  • Establish and govern Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational health metrics.
  • Lead production readiness reviews, disaster recovery testing, resilience assessments, and operational risk mitigation activities.
  • Drive continuous improvement of application stability, service availability, and customer experience.

Incident & Problem Management

  • Lead major incident response and escalation management for critical production issues.
  • Oversee root cause analysis (RCA) processes and ensure corrective actions are implemented and tracked to completion.
  • Drive reduction of recurring incidents through engineering improvements, automation, and proactive monitoring.
  • Provide executive-level communication during significant incidents and service disruptions.

Observability & Automation

  • Establish monitoring, alerting, logging, tracing, and observability standards across supported platforms.
  • Lead implementation of dashboards and operational metrics that provide visibility into service health and customer impact.
  • Drive automation initiatives that reduce manual operational effort, improve recovery times, and increase engineering efficiency.
  • Promote Infrastructure as Code (IaC), CI/CD integration, automated remediation, and self-service operational capabilities.

Cloud & Platform Reliability

  • Partner with Engineering and Infrastructure teams to support cloud-native and hybrid application environments.
  • Ensure applications are designed and operated using resilient, scalable, and supportable architectures.
  • Support modernization initiatives involving Azure cloud services, containers, APIs, microservices, and platform engineering practices.
  • Evaluate vendor platforms and third-party services to ensure reliability and operational readiness.

AI & Modern Operations

  • Drive adoption of AI and Generative AI capabilities to improve incident response, troubleshooting, observability, and operational efficiency.
  • Identify opportunities for intelligent automation, anomaly detection, automated diagnostics, and AI-assisted knowledge management.
  • Promote responsible AI adoption aligned with enterprise security, governance, and risk standards.

People Leadership

  • Recruit, develop, coach, and retain high-performing Site Reliability Engineers, Production Engineers, Automation Engineers, and Observability Engineers.
  • Establish career paths, skill development plans, and succession strategies.
  • Foster a culture of ownership, accountability, innovation, collaboration, and continuous learning.
  • Manage staffing, performance management, compensation recommendations, and organizational development activities.

Risk & Governance

  • Ensure adherence to enterprise risk, cybersecurity, regulatory, and operational control standards.
  • Identify and escalate operational risks impacting critical services or customer experiences.
  • Support audits, regulatory reviews, disaster recovery exercises, and operational governance programs.

Scope of Responsibilities

Leads teams responsible for:

  • Site Reliability Engineering (SRE)
  • Production Support
  • Observability Engineering
  • Incident Management
  • Operational Automation
  • Cloud Reliability
  • Platform Operations

Responsible for reliability and operational health across multiple applications, platforms, cloud services, and vendor-supported solutions.

Supervisory Responsibilities

Typically manages 10-20 direct and indirect reports including SRE Engineers, Production Engineers, Technical Leads, and Engineering Managers.

Education & Experience Required

  • 10+ years of technology experience with application support, infrastructure, cloud, software engineering, or reliability engineering responsibilities.
  • 5+ years of leadership experience managing engineering, operations, or SRE teams.
  • Experience managing production systems supporting critical business functions.
  • Strong knowledge of Site Reliability Engineering principles, including SLOs, observability, automation, incident management, and operational excellence.
  • Experience leading major incident response, root cause analysis, and service restoration efforts.
  • Experience with cloud platforms, distributed systems, APIs, and modern application architectures.
  • Strong communication, analytical, decision-making, and stakeholder management skills.

Preferred Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Technology, or related field.
  • Experience leading SRE or Production Engineering organizations.
  • Experience with Azure cloud technologies and cloud-native architectures.
  • Experience with observability platforms such as Dynatrace, Splunk, Datadog, Grafana, Azure Monitor, or OpenTelemetry.
  • Experience with scripting and automation technologies including PowerShell, Python, Bash, and APIs.
  • Experience with CI/CD, Infrastructure as Code, DevOps, and Platform Engineering practices.
  • Experience implementing operational AI use cases including incident analysis, observability analytics, and automated diagnostics.
  • Financial services or other highly regulated industry experience preferred.

What Great Looks Like

A successful SRE Manager at M&T:

  • Delivers highly available and resilient customer-facing platforms.
  • Uses automation to eliminate operational toil and improve efficiency.
  • Reduces mean time to detect (MTTD) and mean time to restore (MTTR).
  • Establishes strong observability and operational intelligence capabilities.
  • Builds a culture of reliability, accountability, and continuous improvement.
  • Successfully integrates AI-assisted operations and automation into support workflows.
  • develops high-performing teams that balance reliability, speed, risk management, and customer experience.
M&T Bank is committed to fair, competitive, and market-informed pay for our employees. The pay range for this position is $139,700.00 - $232,900.00 Annual (USD). The successful candidate’s particular combination of knowledge, skills, and experience will inform their specific compensation.

Location

Buffalo, New York, United States of America
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
707,525 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Buffalo
$140k – $233k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Buffalo
AI/ML
Agentic Workflows
Machine Learning
DevOps
Azure DevOps
Azure
CI/CD
Kubernetes
Platform Engineering
Azure AKS
Incident Management
Management
Agile
Apply
$97k – $162k per year • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Buffalo
Python
SQL
PowerShell
Bash
AI/ML
Copilot
DevOps
Rest API
Splunk
OpenTelemetry
Datadog
Dynatrace
Prometheus
Azure
CI/CD
Kubernetes
Grafana
Self-Healing
AppDynamics
Azure AKS
Incident Management
Apply
$116k – $194k per year • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Buffalo
Python
SQL
PowerShell
Bash
AI/ML
Copilot
DevOps
Rest API
Splunk
OpenTelemetry
Datadog
Dynatrace
Prometheus
Azure
CI/CD
Kubernetes
Grafana
Self-Healing
AppDynamics
Azure AKS
Incident Management
Apply
$87k – $198k per year • In office • TS/SCI • Full-Time • 8+ years exp • Bachelor's Degree • Chantilly
Python
JavaScript
PHP
TypeScript
SQL
C#
C++
Node JS
Apex
Bash
Lua
Visual Basic
C#
.NET
Apex
MuleSoft
Databases
MS SQL
Frontend
Angular
React.js
DevOps
Rest API
GCP
Azure
CI/CD
Windows Server
AWS
Analytics
Tableau
Informatica
Management
SharePoint
Agile
Apply
$87k – $198k per year • In office • TS/SCI • Full-Time • 5+ years exp • Bachelor's Degree • Aberdeen
Python
Go
JavaScript
PHP
PowerShell
C#
C++
Node JS
Bash
PHP
Drupal
C#
.NET
Frontend
React.js
DevOps
Azure DevOps
Azure
CI/CD
Windows Server
Jenkins
AWS
GitLab
Linux
Cybersecurity
Active Directory
Management
Agile
Apply
$140k – $233k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Buffalo
AI/ML
Agentic Workflows
Machine Learning
DevOps
Azure DevOps
Azure
CI/CD
Kubernetes
Platform Engineering
Azure AKS
Incident Management
Management
Agile
Apply
$44k – $74k per year • In office • Full-Time • 2+ years exp • High School Diploma • Farmington
Apply
$97k – $162k per year • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Buffalo
Python
SQL
PowerShell
Bash
AI/ML
Copilot
DevOps
Rest API
Splunk
OpenTelemetry
Datadog
Dynatrace
Prometheus
Azure
CI/CD
Kubernetes
Grafana
Self-Healing
AppDynamics
Azure AKS
Incident Management
Apply
$116k – $194k per year • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Buffalo
Python
SQL
PowerShell
Bash
AI/ML
Copilot
DevOps
Rest API
Splunk
OpenTelemetry
Datadog
Dynatrace
Prometheus
Azure
CI/CD
Kubernetes
Grafana
Self-Healing
AppDynamics
Azure AKS
Incident Management
Apply
$42k – $72k per year • In office • Full-Time • 1+ year exp • High School Diploma • Washington
Apply
$196k – $265k per year • Remote/Hybrid • Full-Time • 9+ years exp • Bachelor's Degree • Deerfield • Schaumburg • Evanston • Buffalo • Northbrook
Apply
$103k – $172k per year • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Buffalo
Management
Agile
Apply
Design Drafter 1 day ago
$70k – $90k per year • Equity • In office • Full-Time • 3+ years exp • Associate's Degree • Buffalo
Apply
$70k – $95k per year • Equity • In office • Full-Time • Bachelor's Degree • Buffalo
VHDL
MATLAB
MATLAB
Simulink
Apply
$125k – $165k per year • Equity • Remote • Full-Time • 7+ years exp • Bachelor's Degree • Buffalo
Analytics
Microsoft Excel
Apply
See all jobs
This is one of many
707,525 more open roles from verified company boards, updated every day.